Large Language Models, End Term
The input embeddings for the words “learning”, “brings” and “joy” are , , and , respectively. Note that the embeddings are row vectors. The projection matrices are as follows
The following quantities are computed as
Let denote the unnormalized attention score, denote the normalized attention score (ignore the scaling by ) and denote the linear combination of the value vectors for the word.
Enter the value of first element i.e. with index (0,0) of
The input embeddings for the words “learning”, “brings” and “joy” are $h_1 = [1.0, 0.5, 1]$, $h_2 = [1, 0.25, 0]$, and $h_3 = [0.1, 0.1, 0.9]$, respectively. Note that the embeddings are row vectors. The projection matrices are as follows $$W_Q = \begin{bmatrix} 1 & 1 \\ -1 & 1 \\ 0 & 1 \end{bmatrix} \quad W_K = \begin{bmatrix} 1 & 1 \\ 1 & 0 \\ -1 & 1 \end{bmatrix} \quad W_V = \begin{bmatrix} 0 & 0 \\ -1 & -1 \\ 1 & 1 \end{bmatrix}$$ The following quantities are computed as $$Q = HW_Q \quad K = HW_K \quad V = HW_V$$ Let $e_j$ denote the unnormalized attention score, $a_j$ denote the normalized attention score (ignore the scaling by $\sqrt{d_k}$) and $z_j$ denote the linear combination of the value vectors for the $j - th$ word. Enter the value of first element i.e. with index (0,0) of $\frac{\partial a_3}{\partial e_3}$ Assume that we have a large corpus of text. The vocabulary constructed from the text contains 10000 words. Of these, 100 words occurred only once in the entire corpus of text. The parameters of the embedding layer and the output layer of the model are shared. Suppose we create a batch of 256 samples (each sample is a sentence from the corpus). None of these samples contains any of the 100 rare words. Suppose we pre-train the model for one iteration using the batch of samples, then: Suppose we use a pre-trained model for text generation with the given prompt “I am going to”. Which of the following decoding strategies can be used such that the pre-trained model generates same text completion each time it is executed