01
Neural networks
A neural network composes parameterised transformations and nonlinear activation functions. Stacking layers allows the model to construct hierarchical representations.
h₁ = σ(W₁x + b₁) h₂ = σ(W₂h₁ + b₂) y = Wh₂ + b
02
Backpropagation
Backpropagation efficiently applies the chain rule to calculate derivatives of a loss with respect to the parameters of a computational graph. Gradient-based optimizers then use these derivatives to update the model.
θ ← θ - η∇θL
The algorithm itself is mathematically straightforward. The practical difficulty lies in selecting architectures, objectives, data, optimization schedules and regularization strategies that produce useful representations.
03
Convolutional networks
Convolutional neural networks exploit spatial locality and parameter sharing. The same learned filter can detect a pattern at different positions, making convolution well suited to images and spatial signals.
04
Sequence models
Recurrent architectures process sequences through a persistent hidden state. They can model temporal structure but historically faced optimization and parallelization difficulties for very long sequences.
Attention-based architectures changed this landscape by allowing representations at different positions to interact directly.
05
Transformers
Transformers use attention to construct context-dependent representations. Their high degree of parallelism during training helped make very large pretrained models practical.
Attention(Q,K,V) = softmax(QKᵀ / √dₖ)V
Transformers now support language, vision, audio, multimodal systems and many other domains.