Atlas · GenAI 2026
Reward Modeling
Training a model to score outputs by quality/preference, used as the reward signal in RLHF and for best-of-N selection.
conceptPeak: 2023AlignmentAI consensus: 0/3
Prerequisites
- hardRLHF
The reward model is the signal RLHF optimizes against.
It is a supervised ranking/regression model.
Recommended reference
Training language models to follow instructions with human feedback — arXiv 2203.02155