Question 16
Which of the following is true regarding Hard Attention and Soft Attention?
Soft Attention is smooth and differentiable
Variance reduction techniques are used to train Soft Attention models
Soft Attention is computationally cheaper than Hard Attention when the source input is large
The inference (test time) overhead is low in Hard Attention when compared to Soft Attention models