01
The generalization problem
Training performance measures how well a model fits the observed sample. What matters in deployment is performance on future observations drawn from the relevant data-generating process.
A model can therefore become better at fitting the training data while simultaneously becoming worse at predicting unseen data. This is overfitting.
02
Bias and variance
In many regression settings, expected prediction error can be decomposed into components associated with bias, variance and irreducible noise. High bias corresponds to models that are systematically too restrictive; high variance corresponds to models that respond too strongly to the particular sample.
The useful model is therefore not necessarily the most expressive or the simplest one. It is the one whose capacity, data and assumptions produce the best expected behaviour.
03
Regularization and model selection
Regularization constrains the solution space or introduces a preference for particular parameter values. Cross-validation, held-out validation data, early stopping and architectural choices can all play roles in controlling effective model complexity.
04
Distribution shift
Classical generalization theory often assumes training and deployment examples come from the same distribution. Real systems frequently violate this assumption. Changes in user behaviour, environment, sensors, economics or data collection can create distribution shift.