Atlas · GenAI 2026

RLHF

RLHF (Reward Modeling)

conceptPeak: 2023AlignmentAI consensus: 3/3

Prerequisites

  • RLHF is applied AFTER SFT to align the model with preferences — SFT provides the baseline model that RLHF refines

  • RLHF uses PPO (a policy gradient RL algorithm) to optimize the language model against a reward model

Recommended reference

Ouyang et al. (2022) 'Training Language Models to Follow Instructions with Human Feedback' (InstructGPT paper) — the paper that launched the RLHF era

Notes from AI deep research

Anthropic Opus

Ouyang (2022) InstructGPT. PPO + reward model. Zlozony ale kluczowy dla safety-critical apps

OpenAI Deep Research

Kompromisy safety [OA#17]

Google Deep Think

Preferencje na bazie ocen ludzkich [G#63]

Related skills