Quiz Space

Large Language Models · Quiz 1 · 25 Feb 2024 · January 2024 term

Question 6: Consider the following configuration for the Vannila tran…

Question 6

+3 marksNumerical answer

Consider the following configuration for the Vannila transformer architecture with one encoder layer and one decoder layer.

  • Source and target vocabulary size =100= 100
  • maximum sequence length =32= 32
  • length of context window (TT) for both encoder and decoder =32= 32
  • number of heads nh=2n_h = 2
  • dff=4∗dmodeldff = 4 * dmodel
  • dq=dk=dv=dmodelnhdq = dk = dv = \frac{dmodel}{n_h}

Based on the above data, answer the given subquestions.

How many parameters are there in the multi-head attention layer of the encoder (exclude the parameters in the WO matrix used for linear transformation)?

Show answer

Correct answer: 768

Question 6 of 15 in the IIT Madras BS Large Language Models (LLM) Quiz 1 paper sat on 25 Feb 2024, in the January 2024 term (IIT M DEGREE AN2 EXAM QDB2 25 Feb 2024). It carries 3 marks.

More questions from this paper

  1. Q1Consider a vocabulary \mathcal{V} = (A, T, C, G) and the corresponding embedding matrix E = \begin{bmatrix} 1 & 0 \ 0 &…
  2. Q2Consider a vocabulary \mathcal{V} = (A, T, C, G) and the corresponding embedding matrix E = \begin{bmatrix} 1 & 0 \ 0 &…
  3. Q3Consider a vocabulary \mathcal{V} = (A, T, C, G) and the corresponding embedding matrix E = \begin{bmatrix} 1 & 0 \ 0 &…
  4. Q4Consider a vocabulary \mathcal{V} = (A, T, C, G) and the corresponding embedding matrix E = \begin{bmatrix} 1 & 0 \ 0 &…
  5. Q5Consider the following configuration for the Vannila transformer architecture with one encoder layer and one decoder la…
  6. Q7Figure question
  7. Q8Suppose we run the pre-trained GPT model in an autoregressive fashion for generating a text sequence. Then, the stateme…
  8. Q9Suppose we have a dataset for machine translation tasks with thousands of samples. Suppose a team considers training th…
  9. Q10A team trained a Language model using the GPT architecture using a large corpus of text. The model configurations are g…
  10. Q11Choose the correct statements
  11. Q12Figure question
  12. Q13Consider a GPT model used for Causal language modelling. We feed the input sentence “This is a cool idea” to the model …
  13. Q14Consider a GPT model used for Causal language modelling. We feed the input sentence “This is a cool idea” to the model …
  14. Q15The statement that “the Next Sentence Prediction (NSP) task requires the BERT model to run autoregressively given the f…