Figure from the original question paper Consider two data points $\mathbf{x_1} = \begin{bmatrix} a \\ a \end{bmatrix}$ and $\mathbf{x_2} = -1 * \mathbf{x_1}$, where $a > 0$. The data point $\mathbf{x_1}$ belongs to positive class (denoted as 1) and the datapoint $\mathbf{x_2}$ belongs to negative class (denoted by 0). Suppose that the perceptron learning algorithm is used to find the decision boundary that separates these data points with the following rule, $$f(\mathbf{x}) = \begin{cases} 1, & \text{if } \mathbf{w^T x} \ge 0, \\ 0 & \mathbf{w^T x} < 0 \end{cases}$$ The algorithm checks $\mathbf{x_1}$ in the first iteration and $\mathbf{x_2}$ in the second iteration and so on. How many times the weights get updated until convergence (That is, the algorithm classifies both the points correctly)? The weights do not include bias. Assume the weights are initialized to zero Figure from the original question paper