01
What machine learning means
Machine learning is the study of algorithms that improve their performance through experience. The experience is commonly represented as data, although reinforcement learning obtains experience through interaction with an environment.
The key shift is from explicitly specifying every rule to specifying a family of functions, an objective and a procedure for choosing parameters. The system then uses observations to find parameters that produce useful behaviour.
02
The basic learning system
- Data: observations used to estimate useful patterns.
- Representation: the mathematical form in which the model represents inputs and outputs.
- Objective: a numerical measure of how well the model is performing.
- Optimization: a procedure used to search for parameters that improve the objective.
- Evaluation: an independent estimate of how well the learned system generalizes.
03
Empirical risk minimization
Suppose a model fθ maps an input x to a prediction. A loss function L measures the error of that prediction. Training often minimizes the average loss over the observed dataset.
θ* = argminθ (1/n) Σᵢ L(fθ(xᵢ), yᵢ)
This empirical objective is only a proxy for the real goal: low expected loss on future examples. The distinction between training performance and population performance is the source of the generalization problem.
04
Mathematical foundations
Linear algebra describes vectors, matrices, tensors and transformations. Probability describes uncertainty and distributions. Calculus supplies derivatives and gradients. Optimization turns learning into the problem of finding parameters that improve an objective.
These subjects are not optional decoration around machine learning. They are the language in which models, objectives and training procedures are defined.
05
The learning workflow
- Define the prediction or decision problem.
- Collect and inspect representative data.
- Define a model and objective.
- Split data appropriately for development and evaluation.
- Train the model.
- Evaluate on held-out data.
- Perform error analysis.
- Deploy and monitor the system.
- Update the model when assumptions or data distributions change.