Question 1
Teacher forcing for training an transformer model is:
mandatory.
optional.
Teacher forcing for training an transformer model is:
mandatory.
optional.
Choose the correct statements regarding transformer architecture:
Residual connections are practically optional in a transformer model with 10 encoder layers and 10 decoder layers.
Having multiple heads helps in capturing different relationships between input tokens.
One hot encoding is a good choice for position encoding.
Batch normalization can be used instead of layer normalization.
None of these.
Based on the above data, answer the given subquestions.
Assume the model has two layers (N = 2). Calculate the total number of parameters in the model (excluding the embedding layer and output layer). Moreover, no bias was added to the neuron in the FFNN layers.
Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.
The IIT Madras BS Large Language Models (LLM) Quiz 1 paper sat on 27 Oct 2024, in the September 2024 term: 19 questions for 40 marks in 120 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.
| Feature | LLM Quiz 1 27 Oct 2024 at a glance |
|---|---|
| Term | September 2024 term |
| Subject | Large Language Models |
| Course code | BSDA5004 |
| Questions | 19 |
| Marks | 40 |
| Duration | 120 min |
| MCQ | 5 |
| Numerical | 14 |
| Official paper | IIT M DEGREE AN EXAM QDB2 27 Oct 2024 |
| Negative marking | No negative marking. |
| Updated |