siddhant

Knowledge / Generative AI

Transformers

Self-attention, encoder-decoder architectures, positional information, and the foundation of modern language generation.

By Siddhant Krishna · Published 2026-10-06 · Updated 2026-10-06

01

Architecture

The transformer replaces recurrence as the central sequence-processing mechanism with attention-based interactions. This permits substantial parallelization during training.

  • Token embeddings.
  • Positional representation.
  • Multi-head self-attention.
  • Feed-forward transformations.
  • Residual connections.
  • Normalization.
  • Optional cross-attention.

02

Self-Attention

Attention(Q,K,V) = softmax(QKᵀ / √d_k)V

Multiple attention heads can learn different relationships among tokens. Stacking transformer blocks creates progressively richer contextual representations.

References

  1. Vaswani et al. (2017), Advances in Neural Information Processing Systems.
    https://arxiv.org/abs/1706.03762

Related

Contact

Get in Touch

Want to chat? Just shoot me a dm with a direct question on twitter and I'll respond whenever I can. I will ignore all soliciting.