RLHF
Reinforcement Learning from Human Feedback (RLHF) is a post-training method that turns human comparisons between model outputs into a learning signal. Reviewers rank alternative responses; a reward model learns to predict those preferences; and reinforcement learning updates the policy to obtain higher predicted reward, usually while limiting divergence from a reference model. RLHF differs from supervised fine-tuning because the optimization target is learned from comparative judgments rather than copied directly from demonstration answers.
Origin and context
The core preference-learning setup was demonstrated by Christiano and colleagues in 2017 on Atari games and simulated robot control: people chose between short trajectory segments, and those comparisons were used to learn a reward function. In 2022, OpenAI's InstructGPT work applied a related pipeline to language models using demonstrations and ranked responses, while Anthropic independently reported preference modeling and RLHF for helpful and harmless assistants. These papers mark the transition from a general reinforcement-learning technique to a prominent language-model post-training practice; they do not establish a single inventor of every modern implementation.
Why it matters
RLHF matters because many desired assistant behaviors—following an instruction, choosing a useful level of detail, or declining an unsafe request—are difficult to encode as fixed rules. Comparative judgments let developers express such preferences without writing a complete reward function. In the InstructGPT study, a 1.3-billion-parameter model was preferred by evaluators to the much larger GPT-3 baseline on the study's prompt distribution, illustrating how post-training can change perceived usefulness independently of pretraining scale. Anthropic's results provide separate evidence that the approach can coexist with specialized capabilities, although neither study implies that RLHF guarantees broad alignment.
Example
A typical pipeline begins with supervised fine-tuning on curated demonstrations. For each prompt, the current model then produces several candidate responses, which annotators rank. Those comparisons train a reward model, and a reinforcement-learning algorithm updates the assistant against that learned score while a penalty discourages excessive movement away from the reference policy. The result is evaluated by people and by task-specific tests, not by reward alone. For example, preferences can teach a summarization assistant to balance coverage, clarity, and brevity even when no single reference summary is uniquely correct. Exact pipelines vary; PPO is common historically but is not part of the definition.
Maturity and evidence
RLHF merits maturity 4: multiple independent organizations published detailed applications to language-model assistants in 2022, and the method became a well-established option in post-training. The rating does not mean that every current frontier model uses the same pipeline or that RLHF is the only alignment technique. DPO, AI-feedback methods, rule-based rewards, and mixed training recipes can replace or supplement individual stages. A future downgrade would be justified if the term ceased to describe deployed practice and survived mainly as historical shorthand.
Limits and open questions
Human rankings reflect the sampled prompts, annotator population, instructions, and trade-offs chosen by the developer; they are not a neutral measurement of universal human values. The learned reward is also a proxy, so optimizing it can favor responses that score well without being more truthful or robust outside the training distribution. Both InstructGPT and Anthropic therefore evaluate behavior separately and report remaining errors or competing objectives. RLHF can improve measured preference and reduce some observed harms, but it does not prove factual correctness, eliminate reward gaming, or settle whose preferences a system should follow.
Related terms
References
- Deep reinforcement learning from human preferencesarXiv · 2017-06-12 · class A
- Training language models to follow instructions with human feedbackOpenAI / arXiv · 2022-03-04 · class A
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human FeedbackAnthropic / arXiv · 2022-04-12 · class A
Last updated: 2026-08-27
This term is also covered in the Skills Atlas as RLHF skill.
This term is also covered in the Skills Atlas as reinforcement learning skill.