Reasoning Effort and Thinking Budget
Reasoning effort and thinking budget are provider-exposed controls for trading a reasoning model's computational work against latency and cost. Effort is usually a categorical or behavioral signal such as low, medium or high; a thinking budget allocates or caps a number of reasoning tokens. They address the same operational choice but are not exact synonyms, and neither guarantees that a model will use a precise amount of compute or improve every answer.
Origin and context
OpenAI's December 2024 o1 API release introduced `reasoning_effort` as a way to control how long the model thinks. Anthropic announced Claude 3.7 Sonnet with a developer-set thinking budget on 24 February 2025. Google released Gemini 2.5 Flash with a `thinking_budget` parameter on 17 April 2025, explicitly describing a cap the model need not fully consume. This multi-vendor chronology establishes a durable interface category, while also showing why one provider's parameter semantics should not be copied onto another's.
Why it matters
A single default is inefficient when workloads range from extraction to difficult planning. These controls let an application reserve deeper reasoning for requests where evaluations show a benefit and reduce delay or token spend elsewhere. They also make routing policies testable: teams can compare task accuracy, tool-call quality, latency and cost at different settings. The control belongs in product and evaluation design, not only prompting, because supported values, defaults and billing behavior are part of the model API contract.
Example
A support system might use low reasoning effort for intent classification and a higher level for diagnosing an ambiguous account problem. With Gemini 2.5 Flash, the same experiment could set a numeric thinking budget and observe that the model sometimes stops before reaching the cap. Comparing those conditions is valid only within the documented model and API version. Setting `max_tokens` for the entire response is not necessarily a thinking budget, because it may also constrain visible output and tool arguments.
How it differs
Budget Forcing
Budget forcing changes decoding when a model tries to end its reasoning, for example by appending `Wait` or truncating at a chosen point. Reasoning effort and thinking-budget parameters are service-level controls whose internal implementation may be hidden. Similar goals do not make the mechanisms interchangeable.
Maturity and evidence
Maturity is rated 3. The cited releases document related controls from three independent providers. This page compares their interface semantics; it does not define a shared protocol or interchangeable unit of reasoning. Names, supported settings and the relationship between token allocation and effort differ, so a cross-provider comparison must identify the model and release it describes.
Limits and open questions
A larger allowance is not a guaranteed accuracy improvement, and a numeric cap need not be fully consumed. The cited releases document particular models at particular dates, not the current parameter contract for every descendant model. As an evaluation recommendation, compare settings on the same workload and record the model version, observed latency and outcome quality. Do not equate one provider's categorical effort level with another provider's numeric token budget.
Related terms
References
- OpenAI o1 and new tools for developersOpenAI · 2024-12-17 · class A
- Claude's extended thinkingAnthropic · 2025-02-24 · class A
- Start building with Gemini 2.5 FlashGoogle · 2025-04-17 · class A
- s1: Simple test-time scalingMuennighoff et al. / Association for Computational Linguistics · 2025-11 · class A
Last updated: 2026-09-05
This term is also covered in the Skills Atlas as test time compute scaling skill.
This term is also covered in the Skills Atlas as reasoning models skill.
This term is also covered in the Skills Atlas as ai cost optimization skill.