Large Language Models · Quiz 1 · 27 Oct 2024 · September 2024 term
Question 18: The input embeddings for the words “learning”, “brings” …
Question 18
+4 marksNumerical answer
The input embeddings for the words “learning”, “brings” and “joy” are h1=[0.5,0.25,1], h2=[0.1,0.25,0], and h3=[0.1,0.1,0.9], respectively. Note that the embeddings are row vectors. The projection matrices are as follows
WQ=1−10111WK=010101WV=0−110−11
The following quantities are computed as
Q=HWQK=HWKV=HWV
Let ej denote the unnormalized attention score, aj denote the normalized attention score (ignore the scaling by dk) and zj denote the linear combination of the value vectors for the j−th word.
Suppose the gradient vector ∂z3∂L=[1,2], then what is the gradient vector ∂e3∂L? Enter the sum of gradients.
Question 18 of 19 in the IIT Madras BS Large Language Models (LLM) Quiz 1 paper sat on 27 Oct 2024, in the September 2024 term (IIT M DEGREE AN EXAM QDB2 27 Oct 2024). It carries 4 marks.