Glossary · term

Budget Forcing

Budget forcing is a test-time decoding intervention that makes a reasoning model stop at a chosen token budget or continue after it tries to finish. In the s1 method, shorter runs are terminated and longer runs append the token `Wait` when an end-of-thinking delimiter appears, prompting another reasoning segment. It controls generated reasoning length without changing model weights at request time; it is not the general idea of assigning an API thinking budget.

Training2025-01-31Wave 3 · 2025–26Maturity: 3/5

Origin and context

The s1 arXiv preprint was first submitted on 31 January 2025 and the work was later published at EMNLP 2025. Its authors combined supervised fine-tuning on a curated set of 1,000 reasoning examples with budget forcing at inference. Independent ACL 2025 work then evaluated budget forcing alongside outcome- and process-reward methods on mathematical reasoning in 55 languages. That chronology separates the method's origin from later peer-reviewed publication and evaluation and avoids attributing it vaguely to an anonymous community.

Sources: s1, s2, s3

Why it matters

Budget forcing offers a comparatively simple way to study whether extra serial reasoning tokens help a fixed model, without sampling many complete answers or training a new verifier for every request. It also makes the cost-quality trade-off visible: a system can compare accuracy and latency at several forced lengths. The technique is valuable as an experimental control even when it does not improve a production task. Results should be measured against equal-compute alternatives because a longer trace is not free and is not automatically better.

Sources: s1, s2, s3

Example

Suppose a model normally emits an end-of-thinking marker after 2,000 tokens on a difficult problem. A budget-forcing decoder can suppress that marker, append `Wait`, and let the model continue until a 4,000-token budget; a short-budget condition can terminate the trace earlier. The s1 experiment reported an AIME24 change from 50% to 57% for its Qwen2.5-32B-based system under its setup. That number is a scoped experimental result, not an expected gain for other models or tasks.

Sources: s1, s3

How it differs

Reasoning Effort and Thinking Budget

Provider reasoning controls ask a model or service to use a selected effort level or token allowance. Budget forcing directly intervenes in decoding when the model attempts to end its reasoning. The controls can pursue a similar cost-quality trade-off, but their mechanisms and guarantees are not equivalent.

Maturity and evidence

Maturity is rated 3. Budget forcing has a defined originating method, peer-reviewed publication and independent experimental use outside the s1 team. It remains below 4 because evidence is concentrated in reasoning benchmarks, implementations depend on model-specific delimiters and stopping behavior, and equal-compute comparisons do not show a universal advantage over alternatives such as best-of-N or reward-guided selection.

Sources: s2, s3

Limits and open questions

A model can spend the extra tokens repeating itself, following a bad path or reaching a context limit. Appending one token assumes the model learned a useful response to that cue, while forced truncation may cut off an answer. Independent multilingual evaluation found modest, uneven gains and performance comparable to traditional scaling methods under similar inference FLOPs. Teams should report the model, prompt, delimiter, budget, compute accounting and stopping rule rather than treating reasoning length as a quality proxy.

Sources: s1, s2, s3

Related terms

References

Last updated: 2026-09-04

In the Skills Atlas

This term is also covered in the Skills Atlas as test time compute scaling skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as llm decoding strategies skill.