01
Reinforcement Learning from Human Feedback
In the InstructGPT approach, human-written demonstrations are used for supervised fine-tuning and human rankings of outputs are used to train a reward model, followed by reinforcement learning against that reward.
02
Alignment Failure
A model can optimize a proxy objective while failing the underlying human intent. This is a general specification problem: the measurable objective is only an approximation of the desired behavior.