siddhant

Knowledge / Generative AI

Alignment and RLHF

Instruction following, preference learning, reward modeling, and behavioral alignment.

By Siddhant Krishna · Published 2026-10-06 · Updated 2026-10-06

01

Reinforcement Learning from Human Feedback

In the InstructGPT approach, human-written demonstrations are used for supervised fine-tuning and human rankings of outputs are used to train a reward model, followed by reinforcement learning against that reward.

02

Alignment Failure

A model can optimize a proxy objective while failing the underlying human intent. This is a general specification problem: the measurable objective is only an approximation of the desired behavior.

References

  1. Ouyang et al. (2022).
    https://arxiv.org/abs/2203.02155
  2. NIST AI 600-1 (2024; updated 2026).
    https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence
Contact

Get in Touch

Want to chat? Just shoot me a dm with a direct question on twitter and I'll respond whenever I can. I will ignore all soliciting.