Quiz Space

Large Language Models Quiz 2: 24 March 2024 (January 2024 term)

Question 1

+3 marksNumerical answer

Three teams, namely A, B, and C, decided to use a GPT model with two decoder layers with the following configuration

  • length of context window (TT) =1024= 1024
  • number of heads nh=8n_h = 8
  • dmodel=512dmodel = 512
  • dff=4∗dmodeldff = 4 * dmodel
  • dq=dk=dv=dmodelnhdq = dk = dv = \frac{dmodel}{n_h}
  • The weights of the embedding layer and the output layer are shared (tied)

They train the model on the pre-training dataset containing one million words after applying normalization and pre-tokenization( using white space as a delimiter). Moreover, the dataset has no duplicate sentences and no upper-case letters or words. Suppose the team prefers to use the BPE tokenizer to build the vocabulary and subsequently use that to train the model. The base vocabulary contains lowercase alphabets (a to z), digits (0 to 9) and special tokens [unk],[go] and [end].

  • Team AA uses base vocabulary
  • Team BB takes the vocabulary from team AA and does 500 merges
  • Team CC takes the vocabulary from the team BB and does additional 500 merges

Based on the above data answer the given subquestions.

What is the size of the vocabulary built by the Team C?

Question 2

+4 marksOne or more correct options

Three teams, namely A, B, and C, decided to use a GPT model with two decoder layers with the following configuration

  • length of context window (TT) =1024= 1024
  • number of heads nh=8n_h = 8
  • dmodel=512dmodel = 512
  • dff=4∗dmodeldff = 4 * dmodel
  • dq=dk=dv=dmodelnhdq = dk = dv = \frac{dmodel}{n_h}
  • The weights of the embedding layer and the output layer are shared (tied)

They train the model on the pre-training dataset containing one million words after applying normalization and pre-tokenization( using white space as a delimiter). Moreover, the dataset has no duplicate sentences and no upper-case letters or words. Suppose the team prefers to use the BPE tokenizer to build the vocabulary and subsequently use that to train the model. The base vocabulary contains lowercase alphabets (a to z), digits (0 to 9) and special tokens [unk],[go] and [end].

  • Team AA uses base vocabulary
  • Team BB takes the vocabulary from team AA and does 500 merges
  • Team CC takes the vocabulary from the team BB and does additional 500 merges

Based on the above data answer the given subquestions.

Choose the correct statements about the number of parameters in the model (excluding the embedding and output layer parameters)

Select all that apply.

  1. A

    Team A model has less number of parameters than the team B

  2. B

    Team B model has less number of parameters than the team C

  3. C

    Team C model has more number of parameters than the team A

  4. D

    All models have the same number of parameters

Question 3

+4 marksOne or more correct options

Three teams, namely A, B, and C, decided to use a GPT model with two decoder layers with the following configuration

  • length of context window (TT) =1024= 1024
  • number of heads nh=8n_h = 8
  • dmodel=512dmodel = 512
  • dff=4∗dmodeldff = 4 * dmodel
  • dq=dk=dv=dmodelnhdq = dk = dv = \frac{dmodel}{n_h}
  • The weights of the embedding layer and the output layer are shared (tied)

They train the model on the pre-training dataset containing one million words after applying normalization and pre-tokenization( using white space as a delimiter). Moreover, the dataset has no duplicate sentences and no upper-case letters or words. Suppose the team prefers to use the BPE tokenizer to build the vocabulary and subsequently use that to train the model. The base vocabulary contains lowercase alphabets (a to z), digits (0 to 9) and special tokens [unk],[go] and [end].

  • Team AA uses base vocabulary
  • Team BB takes the vocabulary from team AA and does 500 merges
  • Team CC takes the vocabulary from the team BB and does additional 500 merges

Based on the above data answer the given subquestions.

Assume that the word “acrophobia” is not present in the vocabulary, then which of the following tokenizer(s) is(are) guaranteed to tokenize this word into sub-words units (i.e., it does not return [unk] token)

Select all that apply.

  1. A

    The tokenizer used by the team A

  2. B

    the tokenizer used by the team B

  3. C

    the tokenizer used by the team C

  4. D

    None of the given options

11 more questions in this paper

Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.

More on the LLM Quiz 2 24 Mar 2024 paper

The IIT Madras BS Large Language Models (LLM) Quiz 2 paper sat on 24 Mar 2024, in the January 2024 term: 14 questions for 50 marks in 120 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.

FeatureLLM Quiz 2 24 Mar 2024 at a glance
TermJanuary 2024 term
SubjectLarge Language Models
Course codeBSDA5004
Questions14
Marks50
Duration120 min
Numerical1
MSQ9
MCQ4
Official paperIIT M DEGREE AN EXAM QDB2 24 Mar 2024
Negative markingNo negative marking.
Updated

Same Quiz 2, other subjects

More LLM