Deep Learning, Quiz 2
Suppose that a neural network has millions of parameters (weights and biases). A team decides to use an optimization algorithm with a learning rate scheme that is local to each individual parameters in the network. Moreover, the learning rate changes in each iteration based on the magnitude of gradients pertaining to a parameter in the past. Which of the following optimization algorithms satisfy the team’s requirements?
Suppose that a neural network has millions of parameters (weights and biases). A team decides to use an optimization algorithm with a learning rate scheme that is local to each individual parameters in the network. Moreover, the learning rate changes in each iteration based on the magnitude of gradients pertaining to a parameter in the past. Which of the following optimization algorithms satisfy the team’s requirements? Suppose a data set has $N$ samples and each sample has two features $f_1 \in \{0, 1\}$ and $f_2 \in \{0, 1\}$. Assume that $f_1$ is a sparse feature and $f_2$ is a dense feature (that is, $f_1$ value for most of the samples is zero). Further, we apply Stochastic Gradient Descent on this data and plot a contour plot of the loss surface and observe the movement of parameters (i.e., trajectory) across iterations. Let $w_1$ and $w_2$ be the parameters corresponding to $f_1$ and $f_2$ respectively. Assume bias to be zero for this question and the parameters are initialized to zero. In the contour plot of a loss surface, what is the initial movement expected to look like if $w_1$ is plotted on the horizontal axis and $w_2$ is plotted on the vertical axis? Hint: Think of the possible configurations of inputs for the first few iterations Figure from the original question paper