Figure from the passage in the original paper Figure from the original question paper Figure from the passage in the original paper Suppose we transform the data points $X$ using a mapping $\Phi(\cdot)$. Assume there exists a kernel matrix $K$ for the mapping $\Phi(\cdot)$. Moreover, we categorize the model as parametric and non-parametric according to the following definitions. - **Parametric:** The samples in the training set are not necessary for making predictions on a test sample - **non-Parametric:** All the samples in the training set are necessary for making predictions on a test sample Check all that is true about the kernel regression. Figure from the passage in the original paper The team modifies the relation as $y = X^T w + \epsilon$ where $\epsilon \sim \mathcal{N}(0, \sigma^2)$. Suppose that $\sigma^2 = 1$, $d = 3$ and $n = 100$. Assume that $\sum_{i=1}^{n} (w^T x_i - y_i)^2 = 0$ for $w = w^*$. What is the **negative** log-likelihood of the dataset $(X, y)$ at $w^*$? Use logarithm to base $10$.