Deep Learning, Quiz 2
Given a quadratic loss function , where represents the model parameter, you are using the AdaGrad optimizer with Stochastic Gradient Descent to minimize this loss. The learning rate is set to 1, and at the initial iteration (), the parameter has a starting value of .
AdaGrad Update Rule:
Based on the above data, answer the given subquestions.
Given a quadratic loss function $L(w) = w^2$, where $w$ represents the model parameter, you are using the AdaGrad optimizer with Stochastic Gradient Descent to minimize this loss. The learning rate $\eta$ is set to 1, and at the initial iteration ($t=0$), the parameter $w$ has a starting value of $w=2$. **AdaGrad Update Rule:** $$v_t = v_{t-1} + (\nabla w_t)^2$$ $$w_{t+1} = w_t - \frac{\eta}{\sqrt{v_t + \epsilon}} * \nabla w_t$$ $$v_{-1} = 0$$ $$\text{use } \epsilon = 0$$ Based on the above data, answer the given subquestions. Figure from the original question paper Given a quadratic loss function $L(w) = w^2$, where $w$ represents the model parameter, you are using the AdaGrad optimizer with Stochastic Gradient Descent to minimize this loss. The learning rate $\eta$ is set to 1, and at the initial iteration ($t=0$), the parameter $w$ has a starting value of $w=2$. **AdaGrad Update Rule:** $$v_t = v_{t-1} + (\nabla w_t)^2$$ $$w_{t+1} = w_t - \frac{\eta}{\sqrt{v_t + \epsilon}} * \nabla w_t$$ $$v_{-1} = 0$$ $$\text{use } \epsilon = 0$$ Based on the above data, answer the given subquestions. Figure from the original question paper Given a quadratic loss function $L(w) = w^2$, where $w$ represents the model parameter, you are using the AdaGrad optimizer with Stochastic Gradient Descent to minimize this loss. The learning rate $\eta$ is set to 1, and at the initial iteration ($t=0$), the parameter $w$ has a starting value of $w=2$. **AdaGrad Update Rule:** $$v_t = v_{t-1} + (\nabla w_t)^2$$ $$w_{t+1} = w_t - \frac{\eta}{\sqrt{v_t + \epsilon}} * \nabla w_t$$ $$v_{-1} = 0$$ $$\text{use } \epsilon = 0$$ Based on the above data, answer the given subquestions. Figure from the original question paper