Glossary · term

Self-Rewarding Models (SRM)

A self-rewarding model is a language model trained in an iterative loop in which the model also judges candidate responses and turns those judgments into preference or reward signals for its own improvement. The same model family can therefore play both learner and evaluator roles. This is narrower than LLM-as-a-judge, which can evaluate outputs without updating the judging model, and broader than any one preference optimizer.

Training2024-01-18Wave 2 · 2024Maturity: 3/5

Origin and context

Yuan and collaborators submitted Self-Rewarding Language Models in January 2024. Their pipeline generated responses, used an instruction-following rubric to have the model score them, converted comparisons into preference data, and iteratively trained with Direct Preference Optimization. A separate 2025 Findings of ACL paper retained the self-rewarding premise but introduced long reasoning, step-wise judging, and step-wise preference optimization for mathematics, demonstrating independent use of the category beyond the original team.

Sources: s1, s2

Why it matters

Human preference labels are costly and can become a bottleneck when model behavior changes between training rounds. Self-rewarding offers a way to generate fresh supervision at model speed and to improve an evaluator alongside the response policy. The practical attraction is not autonomous self-improvement without limits; it is a reusable data-generation loop. Its quality still depends on the model's rubric interpretation, comparative judgment, sampling diversity, and resistance to reinforcing its own systematic errors.

Sources: s1, s2

Example

For an instruction-following dataset, the current model produces several answers to each prompt and scores them against an explicit rubric. The pipeline retains a preferred and a rejected answer, trains the model on those comparisons, and repeats the cycle with the updated checkpoint. In a process-based variant, the judge evaluates intermediate mathematical steps rather than only the completed answer. A conventional external reward-model pipeline is a counterexample because its reward signal comes from a separately trained evaluator.

Sources: s1, s2

How it differs

LLM-as-a-judge

LLM-as-a-judge names an evaluation role. A self-rewarding loop uses that role to create training signals for the judging model or its successor; an LLM judge used only for benchmarking is not a self-rewarding model.

Reinforcement Fine-Tuning (RFT)

Reinforcement fine-tuning is a broader reward-driven post-training category that can use human or AI feedback. A self-rewarding method is narrower: the language model itself supplies rewards through model-as-judge prompting for its own iterative training.

Maturity and evidence

Maturity is rated 3. The method has a clear primary formulation and an independent peer-reviewed extension with a materially different judging granularity. Evidence remains research-centered, with results tied to selected model families and tasks rather than stable, broadly validated production practice.

Sources: s1, s2

Limits and open questions

The primary study reports only three iterations in one experimental setting, identifies length bias in its judge, and leaves reward hacking as an open question. The independent mathematics extension found that the original approach could be ineffective on mathematical reasoning and evaluated its process-based alternative on selected mathematics tasks. These results do not establish monotonic improvement across domains, so response quality and judge quality should be evaluated separately.

Sources: s1, s2

Related terms

References

Last updated: 2026-09-04

In the Skills Atlas

This term is also covered in the Skills Atlas as reward modeling skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as model training skill.