Glossary · term

Group Relative Policy Optimization (GRPO)

Group Relative Policy Optimization (GRPO) is an online reinforcement-learning algorithm for updating a policy from groups of completions sampled for the same prompt. It computes an advantage by comparing each completion's reward with rewards in its group, then optimizes the policy with a clipped objective and optional reference-policy penalty. GRPO removes PPO's separately trained value or critic model; it does not inherently remove reward functions or reward models.

Training2024-02-05Wave 1 · 2023Maturity: 3/5

Origin and context

DeepSeek introduced GRPO in the DeepSeekMath paper submitted in February 2024, primarily to reduce the memory overhead associated with PPO while training mathematical reasoning. DeepSeek-R1 later used GRPO-family reinforcement learning in a much more visible reasoning-model program. Hugging Face's independent TRL implementation subsequently exposed the algorithm, reward interfaces, and several revised loss variants to practitioners.

Sources: s1, s2, s3

Why it matters

A critic model can be expensive to train and hold in memory alongside the policy and reference model. By estimating relative advantages within each sampled group, GRPO can simplify that part of the reinforcement-learning stack. Its usefulness is broader than mathematics when a task supplies defensible reward signals, but the method's popularity should not be confused with evidence that it is universally cheaper, more stable, or better than PPO.

Sources: s1, s2, s3

Example

For one math prompt, a trainer samples eight candidate solutions, scores each with an answer checker, normalizes those scores within the group, and increases the likelihood of relatively better completions. The same structure can use a learned reward model or a callable reward function. If training merely selects the highest-scoring output without updating a policy, it is best-of-N sampling rather than GRPO.

Sources: s1, s2

Maturity and evidence

Maturity is rated 3. GRPO has a clear originating paper, a prominent later application, maintained independent implementation support, and active research that tests and revises its objective. It remains below 4 because published variants differ materially, reliable outcomes depend on reward and sampling design, and independent work has identified optimization biases in the original formulation.

Sources: s1, s2, s3, s4

Limits and open questions

Group-relative normalization requires multiple completions per prompt and can be uninformative when every completion receives the same reward. Reward quality, group size, clipping, KL settings, and loss normalization all affect training. Independent analysis found a response-length bias in the original objective, while current libraries expose modified formulations. Implementations should report the exact loss and avoid presenting GRPO as synonymous with RLVR or reasoning training generally.

Sources: s2, s4

Related terms

References

Last updated: 2026-09-03

In the Skills Atlas

This term is also covered in the Skills Atlas as reinforcement learning skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as reinforcement learning from verifiable rewards skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as reward modeling skill.