Deep Learning, End Term
Consider a scenario where you have a dataset with overlapping classes (that is instances from different classes share similar or identical feature values), and you decide to train a perceptron model for classification.
Assertion (A): The perceptron model may struggle to classify instances accurately when classes overlap in the feature space.
Reason (R): The perceptron learning algorithm aims to find a linear decision boundary that separates the classes, and in the presence of overlapping classes, it may not be able to capture the underlying patterns effectively.
Select the correct option:
Consider a scenario where you have a dataset with overlapping classes (that is instances from different classes share similar or identical feature values), and you decide to train a perceptron model for classification.\ **Assertion (A):** The perceptron model may struggle to classify instances accurately when classes overlap in the feature space.\ **Reason (R):** The perceptron learning algorithm aims to find a linear decision boundary that separates the classes, and in the presence of overlapping classes, it may not be able to capture the underlying patterns effectively.\ Select the correct option: Consider a feedforward neural network with one hidden layer trained using backpropagation for a binary classification task. The network has the following architecture: - Input layer with 15 neurons - Hidden layer with 25 neurons - Output layer with 1 neuron During the backpropagation process, the derivative of the sigmoid activation function $\sigma(z)$ with respect to its argument $z$ is given by: $$\sigma'(z) = \sigma(z) \cdot (1 - \sigma(z))$$ If the loss function used for binary classification is the binary cross-entropy loss, and the activation fuction at hidden layer and output layer is sigmoid. The output of the neural network is denoted as $\hat{y}$, and the true label is denoted as $y$, what is the expression for $\frac{\partial L}{\partial w_j}$, where $w_j$ represents the weights connecting the $j$th neuron of hidden layer to the output layer? Assume that the output of $j$th neuron of hidden layer is $h_j$ and no biases in the network. In the context of mini-batch gradient descent, if doubling the size of the mini-batch makes your model take twice as many epochs to reach convergence, how does this affect the total number of parameter updates compared to using the original mini-batch size? Assume everything else remains constant.