Question 14
What happens if you increase the number of attention heads in a transformer model while keeping the embedding size constant?
Each head's dimensionality decreases.
The computational complexity of the attention mechanism decreases.
The model learns more diverse relationships in the input sequence.
The model processes longer sequences without truncation.