Question 13
Vanishing gradients in vanilla RNNs are largely caused by:
The continuous addition of bias terms at each timestep which suppresses the gradient signal
Repeated multiplication by Jacobians whose spectral norm is often < 1
The use of unbounded activation functions like ReLU that cause activations to decay over time
The inability of the standard cross-entropy loss function to distinguish between short-term and long-term dependencies