Glossary · term

On-Policy Distillation

On-policy distillation, or OPD, trains a student model on sequences sampled from the student's current policy while a teacher supplies token-level targets or probability feedback on those same sequences. It combines student-visited training contexts with the dense supervision of knowledge distillation. The method is not ordinary self-training: the supervising distribution comes from a teacher, even though the student determines which trajectories are visited.

Training2023-06-23Wave 3 · 2025–26Maturity: 3/5

Origin and context

A June 2023 paper introduced Generalized Knowledge Distillation for autoregressive language models and explicitly framed its student-generated component as on-policy distillation. The work was accepted at ICLR 2024 and allowed mixtures of student- and teacher-generated sequences with configurable divergence losses. In October 2025, Thinking Machines published an independent implementation using student rollouts and teacher log probabilities as dense token-level feedback. A 2026 preprint then examined OPD failure modes and conditions such as compatible teacher–student reasoning patterns and genuinely new teacher capability.

Sources: s1, s2, s3

Why it matters

Off-policy distillation trains on contexts produced by a teacher or fixed dataset. At inference, an autoregressive student instead conditions on its own earlier tokens, including mistakes, and may visit states absent from training. OPD narrows that mismatch by supervising the student's actual rollouts. Compared with a single outcome reward, teacher probabilities can provide feedback at many token positions. The tradeoff is operational: training must generate from the student and evaluate the resulting sequences with a capable teacher.

Sources: s1, s2, s3

Example

A smaller reasoning model receives a batch of math prompts and generates its own solution traces. A larger teacher computes next-token probabilities over each student trace, and the optimizer updates the student to reduce a selected divergence on those visited states. The next batch is sampled from the updated student, so the data distribution changes with training. If the team instead trains only on completed solutions generated once by the teacher, it is off-policy sequence distillation rather than OPD.

Sources: s1, s2

How it differs

Knowledge distillation

Knowledge distillation is the broader transfer of a teacher's behavior or distribution to a student. OPD specifies that training trajectories are sampled from the student's current policy; conventional sequence distillation commonly uses fixed teacher-generated outputs.

Reinforcement Fine-Tuning (RFT)

Both methods can train on student rollouts. RFT normally optimizes a scalar or sequence-level reward, while OPD uses a teacher distribution or token-level targets as the supervisory signal. Implementations may combine them, but they are not synonyms.

Maturity and evidence

Maturity is rated 3. OPD has a peer-reviewed formulation, an independent end-to-end implementation, and a later systematic study of training dynamics. It remains below 4 because recipes, loss choices, teacher access, and long-horizon behavior are unsettled, and evidence is concentrated in selected model families and benchmark tasks.

Sources: s1, s2, s3

Limits and open questions

The teacher must assign useful probability mass on states the student visits; a large capability or reasoning-style mismatch can make feedback ineffective. Reverse-KL variants may be mode-seeking and cannot easily teach tokens outside the student's practical support without a suitable initialization. Student sampling and teacher scoring also consume compute, while apparent benchmark efficiency depends on how rollout, inference, and training costs are counted.

Sources: s1, s2, s3

Related terms

References

Last updated: 2026-09-04

In the Skills Atlas

This term is also covered in the Skills Atlas as knowledge distillation skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as model training skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as reinforcement learning skill.