Question 18
A GPT-style transformer block has an embedding dimension of 768 and uses 12 attention heads. Consider a specific processing task with a very short sequence of only 3 tokens ( ).Compute the total number of weight parameters (ignore biases) in the self-attention module, specifically for the Query, Key, Value, and Output projection matrices. Report your answer in millions, rounded to one decimal place.