Figure from the original question paper You are given a dataset of $10 \times 10$ *grayscale* images. Your goal is to build a 5-class classifier. You have to adopt one of the following two options: - Model A: the input is flattened into a 100-dimensional vector, followed by a fully-connected layer with 5 neurons without bias - Model B: the input is directly given to a convolutional layer with five $10 \times 10$ filters Suppose you make your choice on the basis of the number of parameters in the models, $p_A$ = number of parameters in ModelA, and similarly let $p_B$ = number of parameters in ModelB. In the context of RNNs, what structural feature of LSTMs helps reduce the impact of vanishing gradients?