Quiz Space

Large Language Models Quiz 1: 25 February 2024 (January 2024 term)

Question 1

+3 marksNumerical answer

Consider a vocabulary V=(A,T,C,G)\mathcal{V} = (A, T, C, G) and the corresponding embedding matrix E=[10011−111]E = \begin{bmatrix} 1 & 0 \\ 0 & 1 \\ 1 & -1 \\ 1 & 1 \end{bmatrix}. For example, the 0−th0 - th row of EE is the embedding for 0−th0 - th token of V\mathcal{V} and so on. The following sequence X=(G,C,C,T,A,G)X = (G, C, C, T, A, G) is to be processed by the encoder layer of the transformer architecture by appropriately adding the positional information using the following encoding scheme

P(pos,2i)=sin(pos10(2i/dmodel))P(pos, 2i) = sin\left(\frac{pos}{10^{(2i/dmodel)}}\right)

P(pos,2i+1)=cos(pos10(2i/dmodel))P(pos, 2i+1) = cos\left(\frac{pos}{10^{(2i/dmodel)}}\right)

Assume the index for the position starts from zero. Add positional information to all the tokens in the sequence. Let the resultant sequence be denoted by H=E(X)+PH = E(X) + P, (HH will be used in the subsequent questions,please write it down).Here, E(X)E(X) denotes the embedding for each token in XX

Based on the above data, answer the given subquestions.

Question 2

+2 marksNumerical answer

Consider a vocabulary V=(A,T,C,G)\mathcal{V} = (A, T, C, G) and the corresponding embedding matrix E=[10011−111]E = \begin{bmatrix} 1 & 0 \\ 0 & 1 \\ 1 & -1 \\ 1 & 1 \end{bmatrix}. For example, the 0−th0 - th row of EE is the embedding for 0−th0 - th token of V\mathcal{V} and so on. The following sequence X=(G,C,C,T,A,G)X = (G, C, C, T, A, G) is to be processed by the encoder layer of the transformer architecture by appropriately adding the positional information using the following encoding scheme

P(pos,2i)=sin(pos10(2i/dmodel))P(pos, 2i) = sin\left(\frac{pos}{10^{(2i/dmodel)}}\right)

P(pos,2i+1)=cos(pos10(2i/dmodel))P(pos, 2i+1) = cos\left(\frac{pos}{10^{(2i/dmodel)}}\right)

Assume the index for the position starts from zero. Add positional information to all the tokens in the sequence. Let the resultant sequence be denoted by H=E(X)+PH = E(X) + P, (HH will be used in the subsequent questions,please write it down).Here, E(X)E(X) denotes the embedding for each token in XX

Based on the above data, answer the given subquestions.

Question 3

+3 marksNumerical answer

Consider a vocabulary V=(A,T,C,G)\mathcal{V} = (A, T, C, G) and the corresponding embedding matrix E=[10011−111]E = \begin{bmatrix} 1 & 0 \\ 0 & 1 \\ 1 & -1 \\ 1 & 1 \end{bmatrix}. For example, the 0−th0 - th row of EE is the embedding for 0−th0 - th token of V\mathcal{V} and so on. The following sequence X=(G,C,C,T,A,G)X = (G, C, C, T, A, G) is to be processed by the encoder layer of the transformer architecture by appropriately adding the positional information using the following encoding scheme

P(pos,2i)=sin(pos10(2i/dmodel))P(pos, 2i) = sin\left(\frac{pos}{10^{(2i/dmodel)}}\right)

P(pos,2i+1)=cos(pos10(2i/dmodel))P(pos, 2i+1) = cos\left(\frac{pos}{10^{(2i/dmodel)}}\right)

Assume the index for the position starts from zero. Add positional information to all the tokens in the sequence. Let the resultant sequence be denoted by H=E(X)+PH = E(X) + P, (HH will be used in the subsequent questions,please write it down).Here, E(X)E(X) denotes the embedding for each token in XX

Based on the above data, answer the given subquestions.

The bottommost encoder layer of the transformer contains the following parameter matrices

WQ=[100−1],WK=[1010]WV=[0110]WO=[1111]W_Q = \begin{bmatrix} 1 & 0 \\ 0 & -1 \end{bmatrix}, \quad W_K = \begin{bmatrix} 1 & 0 \\ 1 & 0 \end{bmatrix} \quad W_V = \begin{bmatrix} 0 & 1 \\ 1 & 0 \end{bmatrix} \quad W_O = \begin{bmatrix} 1 & 1 \\ 1 & 1 \end{bmatrix}

given that the embeddings are row vectors, compute the attention score to produce the new representation for the token (TT) at the 3-rd position in the sequence XX.

12 more questions in this paper

Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.

More on the LLM Quiz 1 25 Feb 2024 paper

The IIT Madras BS Large Language Models (LLM) Quiz 1 paper sat on 25 Feb 2024, in the January 2024 term: 15 questions for 50 marks in 120 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.

FeatureLLM Quiz 1 25 Feb 2024 at a glance
TermJanuary 2024 term
SubjectLarge Language Models
Course codeBSDA5004
Questions15
Marks50
Duration120 min
Numerical6
MCQ5
MSQ4
Official paperIIT M DEGREE AN2 EXAM QDB2 25 Feb 2024
Negative markingNo negative marking.
Updated

Same Quiz 1, other subjects

More LLM