Question 3
Which of the following is not true about multi-head cross attention?
Keys and values are taken from the encoder outputs
Queries come from the decoder outputs
All heads attend to identical features due to shared projection matrices
Each head learns to attend to different representation subspaces