Large Language Models, End Term
The input embeddings for the words “learning”, “brings” and “joy” are , , and , respectively. Note that the embeddings are row vectors. The projection matrices are as follows
The following quantities are computed as
Let denote the unnormalized attention score, denote the normalized attention score (ignore the scaling by ) and denote the linear combination of the value vectors for the word.
Enter the value of the first element, i.e., with index (0,0) of the matrix containing the partial derivatives, .
The input embeddings for the words “learning”, “brings” and “joy” are $h_1 = [0.5, 0.25, 1]$, $h_2 = [0.1, 0.25, 0]$, and $h_3 = [0.1, 0.1, 0.9]$, respectively. Note that the embeddings are row vectors. The projection matrices are as follows $$W_Q = \begin{bmatrix} 1 & 1 \\ -1 & 1 \\ 0 & 1 \end{bmatrix} \quad W_K = \begin{bmatrix} 0 & 1 \\ 1 & 0 \\ 0 & 1 \end{bmatrix} \quad W_V = \begin{bmatrix} 0 & 0 \\ -1 & -1 \\ 1 & 1 \end{bmatrix}$$ The following quantities are computed as $$Q = HW_Q \quad K = HW_K \quad V = HW_V$$ Let $e_j$ denote the unnormalized attention score, $a_j$ denote the normalized attention score (ignore the scaling by $\sqrt{d_k}$) and $z_j$ denote the linear combination of the value vectors for the $j - th$ word. **Enter the value of the first element, i.e., with index (0,0) of the matrix containing the partial derivatives, $\frac{\partial a_3}{\partial e_3}$.** Choose the correct statements regarding an encoder-decoder transformer model: Consider the following statements about increasing the size of the vocabulary