01
The supervised learning setup
In supervised learning the training data consist of examples paired with target outputs. For regression the target is commonly continuous; for classification it belongs to one or more discrete categories.
Dataset D = {(x₁,y₁), (x₂,y₂), …, (xₙ,yₙ)}The learner uses these examples to construct a function that predicts y from x for examples it has not seen.
02
Regression
Regression estimates a numerical quantity. Linear regression assumes the prediction can be expressed as a linear combination of input features.
ŷ = θᵀx + b
The parameters can be estimated by minimizing a loss such as mean squared error. Linear models are simple, computationally efficient and highly interpretable, but their representational capacity is limited.
03
Classification
Classification assigns observations to categories. Binary logistic regression models the probability of a class using the logistic function.
P(y=1|x) = σ(θᵀx) σ(z) = 1 / (1 + e⁻ᶻ)
The decision boundary is determined by the model and the chosen threshold. Multiclass problems can use extensions such as softmax regression.
04
Discriminative versus generative models
Discriminative methods model the relationship between inputs and targets directly, often estimating P(y|x) or a decision boundary. Generative methods model how the observed data are generated, typically by representing P(x|y) together with P(y).
The distinction is not simply philosophical. The assumptions of the model affect sample efficiency, robustness, calibration and the kind of data needed for training.
05
Common failure modes
- Label noise can teach the model contradictory mappings.
- Class imbalance can make aggregate accuracy misleading.
- Data leakage can produce unrealistically strong validation results.
- Distribution shift can break a model after deployment.
- Spurious correlations can produce high test performance without reliable causal understanding.