Question 2
In the scaled dot-product attention equation
To reduce the number of trainable parameters in the model
To prevent the dot product values from growing too large, which would push the softmax function into regions with extremely small gradients
To normalize the embedding vectors to have a unit length
Question 3
16 more questions in this paper
Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.
More on the LLM Quiz 1 15 Mar 2026 paper
The IIT Madras BS Large Language Models (LLM) Quiz 1 paper sat on 15 Mar 2026, in the January 2026 term: 19 questions for 40 marks in 120 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.
| Feature | LLM Quiz 1 15 Mar 2026 at a glance |
|---|---|
| Term | January 2026 term |
| Subject | Large Language Models |
| Course code | BSDA5004 |
| Questions | 19 |
| Marks | 40 |
| Duration | 120 min |
| MCQ | 12 |
| MSQ | 3 |
| Numerical | 4 |
| Official paper | Large Language Models 15 Mar 26 |
| Negative marking | No negative marking. |
| Updated |