Number of heads: 2 (each head operates on 2 dimensions)
For both the heads dK=dQ=dV=2
Weight matrices for WQ,WK,WV first head :
WQ1,WK1,WV1=[12121212]
Weight matrices for WQ,WK,WV second head :
WQ2,WK2,WV2=[11011201]
The output from both heads are concatenated, then projected by:
WO=0.500000.50000100001
Scaled Dot-Product Attention (Q, K, V )=softmax(dkQTK)VT
Based on the above data, answer the given subquestions.
Concatenate the outputs from both the attention heads, then apply the output projection matrix Wo to produce the final output of the multi-head attention mechanism. Select the correct result of this operation.
Question 15 of 21 in the IIT Madras BS Large Language Models (LLM) Quiz 1 paper sat on 13 Jul 2025, in the May 2025 term (IIT M DEGREE AN EXAM QDB2 13 July 2025). It carries 2 marks.