01
Linear regression
Linear regression predicts a continuous target from a weighted sum of features. In matrix notation, predictions are Xθ. With squared error, the objective has a closed-form solution under appropriate rank conditions.
θ̂ = (XᵀX)⁻¹Xᵀy
In practice, numerical solvers are preferred to explicitly computing the inverse because they are more stable and can scale better.
02
Logistic regression
Logistic regression maps a linear score through the sigmoid function to produce a probability-like quantity for binary classification. Despite its name, it is a classification model rather than a regression model.
The decision boundary remains linear in the feature space, but the probabilistic output makes the model useful for ranking, threshold selection and uncertainty-aware decisions.
03
Regularization
Ridge: loss + λ||θ||²₂ Lasso: loss + λ||θ||₁
Ridge regularization shrinks parameter magnitudes, while L1 regularization can drive some coefficients exactly to zero. Regularization expresses a preference for simpler parameterisations and can improve generalization.
04
Geometric interpretation
A linear classifier separates feature space using a hyperplane. The orientation of the parameter vector determines the direction of the boundary, while the bias term determines its offset.
This geometric perspective connects linear models to support vector machines, large-margin learning and the representation-learning problem: nonlinear features can make a problem linearly separable in a transformed space.