Question 1
Which statements are true about traditional attention used in sequence-to-sequence models?
The decoder decides which encoder states are important.
Attention weights are computed over all decoder hidden states.
The context vector is a weighted sum of encoder states.
Encoder hidden states are ignored after encoding.