Decoupled Clip and Dynamic sAmpling Policy Optimization (DAPO)
Decoupled Clip and Dynamic sAmpling Policy Optimization (DAPO) is a reinforcement-learning algorithm and training recipe for language-model reasoning. Building on group-relative policy optimization, it combines asymmetric clipping, dynamic resampling of prompts, token-level policy-gradient loss and soft penalties for overlong responses. The originating work also released code, data and a trained model around the recipe. DAPO therefore names both a specific set of optimization changes and its reference system, not every open-source reasoning-training pipeline.
Origin and context
The DAPO preprint appeared on 18 March 2025 and was revised on 20 May. The ByteDance Seed publication page dates the accompanying open system to 20 May and lists the work at NeurIPS 2025. The authors presented four techniques intended to stabilize large-scale reinforcement learning from verifiable rewards. A September 2025 independent arXiv preprint compared DAPO with GRPO and proposed alternative clipping and reward-standardization choices, showing that the name had become a reproducible research baseline rather than only a release label.
Why it matters
Reasoning-model reinforcement learning can waste batches when every sampled answer for a prompt gets the same reward, clip useful updates or let very long responses dominate optimization. DAPO packages interventions for those concrete failure modes and provides public artifacts for studying them at scale. That makes it useful to researchers comparing RLVR recipes and to engineers who need to specify more than “we used GRPO.” The contribution is not just a new acronym: it exposes choices about sampling, clipping, loss aggregation and length handling that materially change training behavior.
Example
During math training, a prompt whose sampled completions are all correct or all wrong supplies no within-group reward variation. DAPO's dynamic sampling can skip that group and draw another prompt, while Clip-Higher allows a larger upper ratio bound to preserve exploration and token-level aggregation changes how long responses contribute to the update. This is DAPO only when the specified recipe is used. Filtering zero-variance groups by itself is one technique, not sufficient evidence that an entire training run implements DAPO.
How it differs
Group Relative Policy Optimization (GRPO)
GRPO is the broader group-relative policy-optimization baseline that estimates advantages without a separate critic. DAPO modifies that family with a named collection of clipping, sampling, loss and length-control choices. Results for DAPO should not be attributed to GRPO generally, and a GRPO trainer does not automatically implement DAPO.
Maturity and evidence
Maturity is rated 3. DAPO has a public algorithm, released artifacts, a NeurIPS 2025 publication and an independent comparative preprint examining its mechanisms. The base score of 2 is therefore stale. A score of 4 would require broader evidence of sustained adoption across independent production or research stacks and more stable agreement on which components drive gains; current comparative evidence continues to revise those choices.
Limits and open questions
The four components interact, so a headline result cannot identify one causal improvement without ablations. The reported 50-point AIME 2024 result is tied to the authors' Qwen2.5-32B setup, dataset, compute and evaluation. An independent comparative preprint argues that dynamic sampling reduces sampling efficiency and reports a different multi-component recipe that outperforms DAPO in some tested settings. Implementers should record the exact code revision, clipping bounds, group size, reward rules and loss aggregation rather than using DAPO as a loose synonym for open RLVR.
Related terms
References
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleByteDance Seed and Tsinghua AIR / arXiv · 2025-03-18 · class A
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleByteDance Seed · 2025-05-20 · class A
- DCPO: Dynamic Clipping Policy OptimizationIndependent researchers / arXiv · 2025-09-02 · class A
Last updated: 2026-09-04
This term is also covered in the Skills Atlas as reinforcement learning skill.
This term is also covered in the Skills Atlas as reinforcement learning from verifiable rewards skill.