Atlas · GenAI 2026

Reward Modeling

Training a model to score outputs by quality/preference, used as the reward signal in RLHF and for best-of-N selection.

conceptPeak: 2023AlignmentAI consensus: 0/3

Prerequisites

Recommended reference

Training language models to follow instructions with human feedback — arXiv 2203.02155

Notes from AI deep research