Machine Learning Techniques, End Term
Imagine a dataset characterized by two features, Feature 1 and Feature 2, demonstrating a perfect negative correlation of -1. When applying k-means clustering with k = 3 to this dataset, what is the most likely arrangement of cluster centers that minimizes the within-cluster sum of squares (WCSS)?
Imagine a dataset characterized by two features, Feature 1 and Feature 2, demonstrating a perfect negative correlation of -1. When applying k-means clustering with k = 3 to this dataset, what is the most likely arrangement of cluster centers that minimizes the within-cluster sum of squares (WCSS)? Figure from the original question paper Consider a logistic regression model for a binary classification problem with two features $x_1$ and $x_2$. The feature vector is $\begin{bmatrix} x_1 \\ x_2 \end{bmatrix}$ and labels lie in $\{0, 1\}$. The threshold for inference is 0.5. The dummy feature and the weight corresponding to it can be ignored for this problem. Let $x_1$ be the horizontal axis and $x_2$ be the vertical axis. You are given two feature vectors: $$\mathbf{x_1} = \begin{bmatrix} 1 \\ \sqrt{3} \end{bmatrix}, \mathbf{x_2} = \begin{bmatrix} -1 \\ \sqrt{3} \end{bmatrix}$$ The weight vector makes an angle of $\theta$ with the positive $x_1$ axis (horizontal). Each $\theta$ corresponds to a different classifier. For what range of values of $\theta$ are both $\mathbf{x_1}$ and $\mathbf{x_2}$ predicted to belong to class-1? Hints: - To draw the weight vector $\mathbf{w} = \begin{bmatrix} w_1 \\ w_2 \end{bmatrix}$, plot the point $(w_1, w_2)$ and draw an arrow starting at the origin to this point. - $\tan(60^\circ) = \sqrt{3}$