01
Architecture
The transformer replaces recurrence as the central sequence-processing mechanism with attention-based interactions. This permits substantial parallelization during training.
- Token embeddings.
- Positional representation.
- Multi-head self-attention.
- Feed-forward transformations.
- Residual connections.
- Normalization.
- Optional cross-attention.
02
Self-Attention
Attention(Q,K,V) = softmax(QKᵀ / √d_k)V
Multiple attention heads can learn different relationships among tokens. Stacking transformer blocks creates progressively richer contextual representations.