Question 6
Consider an attention mechanism defined as:
Attention(Q, K, V) = softmax(QKT) V
Which of the following is the most likely consequence of removing the scaling factor in this attention computation?
The attention scores become smaller, leading to uniform attention weights.
The dot product values grow large, causing the softmax to produce extremely peaked distributions and unstable gradients.
The model becomes invariant to the dimensionality of key vectors.
The attention mechanism works as usual.