01
Language Modeling
Autoregressive language models optimize the probability of the next token conditioned on previous context.
L(θ) = -Σ_t log P_θ(x_t | x_{<t})02
Scaling
GPT-3 provided an influential empirical demonstration that increasing language-model scale could substantially improve few-shot performance. Modern systems continue to explore scaling across parameters, data, training computation, inference-time computation, and context.