01
Why unsupervised learning matters
Labelled data can be expensive, ambiguous or unavailable. Unsupervised learning instead asks what structure can be discovered from the observations themselves.
The discovered structure may be clusters, low-dimensional representations, probability distributions or features useful for later tasks.
02
K-means clustering
K-means partitions observations into K groups by alternating between assigning each point to its nearest centroid and recomputing the centroids.
minimize Σᵢ ||xᵢ - μ_{cᵢ}||²The objective is non-convex with respect to the assignments and centroids jointly, so the algorithm can converge to different local optima depending on initialization.
03
Mixture models
A Gaussian mixture model assumes observations are generated by a mixture of Gaussian components. Each point has a probability of belonging to every component rather than a hard assignment.
The expectation-maximization algorithm alternates between estimating latent assignment probabilities and updating model parameters.
04
Principal component analysis
PCA finds orthogonal directions that explain as much variance as possible. It can be derived through eigenvectors of the covariance matrix or through singular value decomposition.
PCA is useful for compression, visualization and noise reduction, but the directions of maximum variance are not automatically the directions most useful for prediction.