Question 2
To reduce the computational cost of the softmax layer during training.
To act as a regularizer similar to Dropout.
Question 3
Why is the standard GPT architecture (Decoder-only with causal masking) generally unsuitable for the Masked Language Modeling (MLM) objective as implemented in BERT?
GPT models are too small to learn bidirectional contexts.
The causal mask in GPT prevents the model from attending to future tokens, making it impossible to use right-side context to predict a masked token.
GPT does not have positional embeddings, which are required for MLM.
GPT uses ReLU activation, while BERT uses GELU, which is required for MLM.
18 more questions in this paper
Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.
More on the LLM Quiz 2 12 Apr 2026 paper
The IIT Madras BS Large Language Models (LLM) Quiz 2 paper sat on 12 Apr 2026, in the January 2026 term: 21 questions for 50 marks in 120 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.
| Feature | LLM Quiz 2 12 Apr 2026 at a glance |
|---|---|
| Term | January 2026 term |
| Subject | Large Language Models |
| Course code | BSDA5004 |
| Questions | 21 |
| Marks | 50 |
| Duration | 120 min |
| MCQ | 10 |
| MSQ | 6 |
| Numerical | 5 |
| Official paper | Large Language Models 07 Apr 26 |
| Negative marking | No negative marking. |
| Updated |