Question 9
Given the vectorized self-attention calculation
The softmax operation is applied to the entire matrix at once (globally), not row-wise.
Given the vectorized self-attention calculation
The softmax operation is applied to the entire matrix at once (globally), not row-wise.
Correct answers
Question 9 of 19 in the IIT Madras BS Large Language Models (LLM) Quiz 1 paper sat on 15 Mar 2026, in the January 2026 term (Large Language Models 15 Mar 26). It carries 3 marks.