Glossary · term

Reinforcement Learning with Verifiable Rewards (RLVR)

Reinforcement Learning with Verifiable Rewards (RLVR) is a post-training method in which a model receives rewards from checks that can be computed automatically, such as matching a known answer, satisfying a formal constraint, or passing executable tests. It retains reinforcement-learning optimization but replaces, for selected tasks, a learned human-preference reward model with a verifier. RLVR is therefore best suited to domains where success can be tested reliably; it is not a synonym for all reinforcement learning used in reasoning models.

Training2024-11-22Wave 1 · 2023Maturity: 3/5

Origin and context

The Tülu 3 report, submitted by Ai2 researchers in November 2024, explicitly introduced the name Reinforcement Learning with Verifiable Rewards and included it in an open post-training recipe. Its experiments used tasks with checkable outcomes alongside supervised fine-tuning and DPO. In January 2025, DeepSeek-R1 independently showed that large-scale reinforcement learning without human-labeled reasoning traces could improve performance on verifiable mathematics, coding, and STEM tasks. DeepSeek used its own multi-stage recipe and did not make Tülu 3's exact implementation universal; together the works document rapid cross-organization adoption of the underlying approach.

Sources: s1, s2

Why it matters

RLVR changes the economics of feedback. A correct-answer checker or test suite can score far more samples than human reviewers can rank, making it possible to explore many candidate solutions and repeatedly update the policy. This is especially useful when the reasoning path is open-ended but the final outcome is testable. Tülu 3 reported targeted gains on its verifiable tasks, while DeepSeek-R1 reported strong results after reinforcement learning without labeled reasoning trajectories. ICLR 2026 research also found evidence that answer-based rewards can improve both final answers and intermediate reasoning on studied math and coding settings. These results are task-specific, not proof of general reasoning.

Sources: s1, s2, s3

Example

For a mathematics prompt, a model samples several solutions. A parser extracts each final answer, and a verifier compares it with the known result; format or constraint checks may provide additional rewards. For coding, a sandbox can run tests instead. The training algorithm increases the probability of responses that pass the verifier, often while controlling update size relative to a reference policy. Unlike RLHF, no person needs to rank every sampled pair. Unlike supervised fine-tuning, the model is not required to imitate a provided reasoning trace. The design quality of the task, parser, and held-out evaluation remains part of the system.

Sources: s1, s2

Maturity and evidence

RLVR merits maturity 3. The term has a clear published definition, reproducible open implementations, independent large-scale use, and an expanding research literature. However, algorithms, reward designs, training-stability practices, and claims about what capabilities are learned remain unsettled. It should be described as an established research and engineering method rather than a universal post-training standard. Evidence of robust gains across more open-ended domains, model families, and held-out evaluations would support a higher rating.

Sources: s1, s2, s3, s4

Limits and open questions

Verifiability applies to the checker, not automatically to the quality of the underlying goal. A weak verifier can accept shortcuts, malformed proofs, modified tests, or outputs that satisfy a narrow criterion while missing the intended task. A 2026 preprint on inductive reasoning reports RLVR-trained models exploiting false positives in an extensional verifier instead of learning the requested general rules. Even when the checker is sound, rewards based only on final answers may leave ambiguity about why performance improved or how well it transfers. Robust use therefore needs sandboxing, hidden tests, adversarial validation, and evaluations the policy cannot directly optimize.

Sources: s3, s4

Related terms

References

Last updated: 2026-08-27

In the Skills Atlas

This term is also covered in the Skills Atlas as reinforcement learning skill.