Deep Learning, Quiz 2
The statement that the bias correction in ADAM optimizer is an absolute requirement for it to converge to a local minimum is
The statement that the bias correction in ADAM optimizer is an absolute requirement for it to converge to a local minimum is Team $A$ constructs a dataset $\mathcal{D} = \{X, y\}$ that contains $N$ samples. Each sample $x_i \in \mathbb{R}$ is uniformly sampled from the function $y = f(x) = 1 + c_1x + c_2x^2 + \cdots + c_px^p$ in the interval $-1 \le x \le 1$, with $p < 80$. The dataset is then given to Team $B$ with the information that the samples were *not* corrupted by noise. They split the dataset into a training set and a test (validation) set. Team $B$ assumes a polynomial function $g(x) = k_0 + k_1x + k_2x^2 + \cdots + k_dx^d$ of unknown degree $d$. Therefore, they decided to vary the degree of the polynomial from 1 to 100. For instance, setting $d = 2$ gives the polynomial $g(x) = k_0 + k_1x_1 + k_2x^2$ and the parameters are estimated using the training set. They measure the mean squared error for each setting. Select the true statement(s). There exists a polynomial degree $d \in \{1, 100\}$ for which Fully connected network diagram with 3 hidden layers (a1/h1, a2/h2, a3/h3), output layer O and weights W1 to W4, with ReLU definition, h2 value and gradient question