Question 1
Transformers process input tokens:
One at a time (sequentially)
In reverse order
All at once (in parallel)
Only after seeing the full input
Transformers process input tokens:
One at a time (sequentially)
In reverse order
All at once (in parallel)
Only after seeing the full input
What is the purpose of the softmax function in the attention mechanism?
Normalize attention scores to a probability distribution
Add non-linearity to the model
To predict the correct class for the loss function
Remove redundant features from the input
What is the main difference between GPT and BERT pre-training objectives?
GPT uses Masked Language Modeling, BERT uses Causal Language Modeling
GPT uses Causal Language Modeling, BERT uses Masked Language Modeling
Both use Masked Language Modeling
Both use Causal Language Modeling
Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.
The IIT Madras BS Large Language Models (LLM) Quiz 1 paper sat on 13 Jul 2025, in the May 2025 term: 21 questions for 50 marks in 120 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.
| Feature | LLM Quiz 1 13 Jul 2025 at a glance |
|---|---|
| Term | May 2025 term |
| Subject | Large Language Models |
| Course code | BSDA5004 |
| Questions | 21 |
| Marks | 50 |
| Duration | 120 min |
| MCQ | 8 |
| MSQ | 4 |
| Numerical | 9 |
| Official paper | IIT M DEGREE AN EXAM QDB2 13 July 2025 |
| Negative marking | No negative marking. |
| Updated |