01
Neural networks
A neural network is a parameterised composition of functions. Each layer transforms an input representation and applies a nonlinear activation. Stacking layers allows the network to represent increasingly complex functions.
h₁ = σ(W₁x + b₁) h₂ = σ(W₂h₁ + b₂) y = Wh₂ + b
02
Backpropagation
Backpropagation applies the chain rule of calculus to compute gradients of a loss function with respect to all model parameters. It allows the network to be trained efficiently even when it contains millions or billions of parameters.
Gradient descent then updates the parameters in a direction intended to reduce the objective. Modern optimisers add mechanisms such as adaptive learning rates and momentum.
03
Convolutional neural networks
Convolutional networks exploit local structure and weight sharing. These properties make them particularly effective for visual data, where nearby pixels are correlated and the same feature detector can be useful in multiple locations.
04
Transformers
Transformers replaced recurrence as the central mechanism for many sequence models. Attention allows each position to combine information from other positions according to learned relevance weights.
Attention(Q,K,V) = softmax(QKᵀ / √dₖ)V
The ability to parallelise training efficiently contributed to the scalability of Transformer-based models across language, vision, audio and multimodal tasks.
05
Scaling
Modern deep learning systems can scale across model parameters, training data and compute. Scaling has produced capabilities that were difficult to obtain with smaller systems, but it also increases infrastructure cost, data requirements and the need for careful evaluation.