LLM Quiz 1 19 Jul 2026 — Question 7
Show answer
Correct answer: 256
Question 7 of 16 in the IIT Madras BS Large Language Models (LLM) Quiz 1 paper sat on 19 Jul 2026, in the May 2026 term (Large Language Models 16 Jul 26). It carries 4 marks.
More questions from this paper
- Which statements are true about traditional attention used in sequence-to-sequence models?
- The query, key, and value projection matrices are given by: Based on the above data, answer the given subquestions.
- The query, key, and value projection matrices are given by: Based on the above data, answer the given subquestions. Cho…
- A Transformer model processes a sequence containing 6 tokens using Multi-Head Attention. The model initially uses 4 att…
- Figure question
- Figure question
- Consider the following statements regarding the use of Teacher Forcing while training autoregressive sequence-to-sequen…
- Based on the above data, answer the given subquestions.
- Choose the option corresponding to the sentence with the highest joint probability under the given language model.
- Using the given probability tree calculate the conditional probability of predicting "apples" given "like" [i.e P(apple…
- Which of the following are components found within a GPT decoder layer?
- The table below represents the conditional probability distribution over the vocabulary. Each column corresponds to a d…
- The table below represents the conditional probability distribution over the vocabulary. Each column corresponds to a d…
- The table below represents the conditional probability distribution over the vocabulary. Each column corresponds to a d…
- The table below represents the conditional probability distribution over the vocabulary. Each column corresponds to a d…