siddhant

Knowledge / Machine Learning

Supervised Learning

Learning a mapping from inputs to labelled targets.

By Siddhant Krishna · Published 2026-10-06 · Updated 2026-10-06

01

The supervised learning setup

In supervised learning the training data consist of examples paired with target outputs. For regression the target is commonly continuous; for classification it belongs to one or more discrete categories.

Dataset D = {(x₁,y₁), (x₂,y₂), …, (xₙ,yₙ)}

The learner uses these examples to construct a function that predicts y from x for examples it has not seen.

02

Regression

Regression estimates a numerical quantity. Linear regression assumes the prediction can be expressed as a linear combination of input features.

ŷ = θᵀx + b

The parameters can be estimated by minimizing a loss such as mean squared error. Linear models are simple, computationally efficient and highly interpretable, but their representational capacity is limited.

03

Classification

Classification assigns observations to categories. Binary logistic regression models the probability of a class using the logistic function.

P(y=1|x) = σ(θᵀx)
σ(z) = 1 / (1 + e⁻ᶻ)

The decision boundary is determined by the model and the chosen threshold. Multiclass problems can use extensions such as softmax regression.

04

Discriminative versus generative models

Discriminative methods model the relationship between inputs and targets directly, often estimating P(y|x) or a decision boundary. Generative methods model how the observed data are generated, typically by representing P(x|y) together with P(y).

The distinction is not simply philosophical. The assumptions of the model affect sample efficiency, robustness, calibration and the kind of data needed for training.

05

Common failure modes

  • Label noise can teach the model contradictory mappings.
  • Class imbalance can make aggregate accuracy misleading.
  • Data leakage can produce unrealistically strong validation results.
  • Distribution shift can break a model after deployment.
  • Spurious correlations can produce high test performance without reliable causal understanding.

References

  1. Stanford University, CS229 Machine Learning. Course materials covering supervised and unsupervised learning, learning theory, regularization, SVMs and reinforcement learning.
    https://cs229.stanford.edu/
  2. Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani and Jonathan Taylor. An Introduction to Statistical Learning.
    https://www.statlearning.com/

Related

Contact

Get in Touch

Want to chat? Just shoot me a dm with a direct question on twitter and I'll respond whenever I can. I will ignore all soliciting.