siddhant

Knowledge / Generative AI

Large Language Models

Token prediction, scaling, pretraining, contextual representation, decoding, and emergent task capabilities.

By Siddhant Krishna · Published 2026-10-06 · Updated 2026-10-06

01

Language Modeling

Autoregressive language models optimize the probability of the next token conditioned on previous context.

L(θ) = -Σ_t log P_θ(x_t | x_{<t})

02

Scaling

GPT-3 provided an influential empirical demonstration that increasing language-model scale could substantially improve few-shot performance. Modern systems continue to explore scaling across parameters, data, training computation, inference-time computation, and context.

References

  1. Vaswani et al. (2017), Advances in Neural Information Processing Systems.
    https://arxiv.org/abs/1706.03762
  2. Brown et al. (2020), Advances in Neural Information Processing Systems.
    https://arxiv.org/abs/2005.14165

Related

Contact

Get in Touch

Want to chat? Just shoot me a dm with a direct question on twitter and I'll respond whenever I can. I will ignore all soliciting.