Glossary · term

Process Reward Model (PRM)

A process reward model, or PRM, is a learned model that scores intermediate steps in a multi-step solution or trajectory. It can provide a score after each step, helping a system rank candidate solutions or supply a training signal. Process supervision is the broader labeling or training regime that evaluates intermediate reasoning; a PRM is one model trained to approximate that feedback. An outcome reward model instead scores the final result.

Training2022-11-25Wave 2 · 2024Maturity: 3/5

Origin and context

Uesato and collaborators reported a 2022 comparison of process- and outcome-based feedback on GSM8K. In 2023, Lightman and collaborators found process supervision stronger than outcome supervision in their MATH experiments and released PRM800K, containing step-level human labels used for their reward model. Math-Shepherd then demonstrated an independently developed PRM trained with automatically constructed process supervision and applied it to reranking and reinforcement learning.

Sources: s1, s2, s3

Why it matters

A final answer can be correct despite invalid reasoning, or wrong after several useful steps. Step-level scores expose a finer signal than a single terminal verdict. They can help select among sampled solutions, identify where a trajectory first goes off course, or shape training toward better intermediate work. That extra granularity costs annotation or synthetic-labeling effort and does not make the learned judge infallible.

Sources: s1, s2, s3

Example

For a math problem, a system generates several worked solutions, divides each into steps, and asks a PRM to score the progression after every step. It aggregates those scores to rerank complete solutions before returning one. Developers compare the ranking with held-out expert labels and final-answer checks. A high PRM score is treated as model evidence, not as a proof that every step is valid.

Sources: s2, s3

How it differs

Outcome Reward Model (ORM)

A PRM evaluates intermediate steps; an outcome reward model evaluates the final result or completed trajectory. They are sibling approaches, not duplicate names. Outcome feedback is often cheaper, while process feedback can reveal reasoning errors that a correct endpoint conceals.

RLHF

RLHF is a broader alignment workflow that can use learned rewards derived from human preferences. A PRM specifies where a reward is assigned within a multi-step trajectory. PRMs can also be used only for verification or reranking, without an RLHF training loop.

Maturity and evidence

Maturity is rated 3. The concept has clear primary comparisons, a large released human-label dataset, and an independent peer-reviewed implementation. Results are still concentrated in mathematical reasoning, and training labels, score aggregation, and transfer behavior vary across systems.

Sources: s1, s2, s3

Limits and open questions

A PRM can learn annotator shortcuts, favor familiar solution styles, or assign locally plausible scores to a globally flawed argument. Automatically generated step labels may scale supervision while importing errors from the labeling procedure. Aggregating step scores can change rankings, and performance on math does not establish reliability in medicine, law, or open-ended agent work. Teams should evaluate calibration, adversarial robustness, domain transfer, and disagreement with qualified reviewers before using PRM scores in consequential decisions.

Sources: s1, s2, s3

Related terms

References

Last updated: 2026-09-03

In the Skills Atlas

This term is also covered in the Skills Atlas as reward modeling skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as RLHF skill.