Question 2
A GPT model is trained using causal language modeling. During training, for a sequence of T = 4 tokens, which of the following correctly represents the attention mask matrix applied to the attention logits?
A GPT model is trained using causal language modeling. During training, for a sequence of T = 4 tokens, which of the following correctly represents the attention mask matrix applied to the attention logits?
Correct answer
Question 2 of 17 in the IIT Madras BS Large Language Models (LLM) End Term paper sat on 31 Aug 2025, in the May 2025 term (IIT M DEGREE AN EXAM QDB3 31 Aug 2025). It carries 2 marks.