Atlas · GenAI 2026
RLHF
RLHF (Reward Modeling)
conceptPeak: 2023AlignmentAI consensus: 3/3
Prerequisites
- hardLLM Fine-Tuning
RLHF is applied AFTER SFT to align the model with preferences — SFT provides the baseline model that RLHF refines
- mediumReinforcement Learning
RLHF uses PPO (a policy gradient RL algorithm) to optimize the language model against a reward model
Recommended reference
Ouyang et al. (2022) 'Training Language Models to Follow Instructions with Human Feedback' (InstructGPT paper) — the paper that launched the RLHF era
Notes from AI deep research
Anthropic Opus
Ouyang (2022) InstructGPT. PPO + reward model. Zlozony ale kluczowy dla safety-critical apps
OpenAI Deep Research
Kompromisy safety [OA#17]
Google Deep Think
Preferencje na bazie ocen ludzkich [G#63]
Related skills
- → is subcategory of: LLM Fine-Tuning(3/3)
- ← is an instance of: Direct Preference Optimization(1/3)