Glossary · term

Reasoning models

Reasoning models are language models trained and deployed to devote intermediate computation to multi-step problems before returning a final answer. They may refine a chain of thought, check intermediate work, backtrack, or try another strategy. The label describes a model category and intended behavior, not a guarantee of logically valid reasoning and not one mandatory training recipe. OpenAI o1 and DeepSeek-R1 are two independently documented examples.

Training2024-09-12Wave 1 · 2023Maturity: 3/5

Origin and context

The current category became visible with OpenAI's public release of o1-preview on 12 September 2024. OpenAI described a model trained with large-scale reinforcement learning to improve its chain of thought and reported gains from both train-time reinforcement learning and additional thinking time at test time. On 22 January 2025, DeepSeek-AI released the first version of its R1 paper, calling R1-Zero and R1 its first-generation reasoning models. R1-Zero used reinforcement learning without supervised fine-tuning as a preliminary stage, while R1 added cold-start data and multi-stage training. Together, the sources support a cross-organization category without proving a single originator of the phrase.

Sources: s1, s2

Why it matters

Reasoning models make additional computation available for tasks where an immediate completion is often insufficient, including mathematics, coding, scientific questions, and structured planning. OpenAI describes o1 learning to identify mistakes, simplify difficult steps, and switch approaches. DeepSeek reports self-reflection, verification, and dynamic strategy adaptation emerging under reinforcement learning. The practical change is not that every response becomes more reliable; it is that developers can choose a model family designed to spend more effort on problems that benefit from decomposition and checking, accepting additional cost and latency where justified.

Sources: s1, s2

Example

For a competition-math problem, a conventional chat model may produce a direct solution in one pass. A reasoning model may instead explore candidate derivations, notice that an assumption fails, backtrack, and attempt a different route before presenting its answer. The intermediate process may remain internal or be summarized by the product. A longer process is therefore an implementation behavior, not evidence by itself that the final result is correct, even when the intermediate trace sounds fluent; the answer still needs task-appropriate verification.

Sources: s1, s2

How it differs

Test-time compute

Reasoning models are a model category. Test-time compute is the amount and allocation of inference work after a prompt arrives. A reasoning model can use a larger or smaller thinking budget, while test-time scaling can also generate, search, or rerank outputs from other models.

Reinforcement Learning with Verifiable Rewards (RLVR)

RLVR is a training method based on rewards that can be checked automatically, especially for domains such as mathematics and code. DeepSeek-R1 shows that reinforcement learning on verifiable tasks can develop reasoning behaviors, but a reasoning model may use multi-stage training or other methods. The category and the recipe are not synonyms.

Maturity and evidence

Maturity is rated 3. The category is documented by at least two independent model developers and is linked to concrete training and inference practices, so it is more than a single-product label. It is not rated 4 because there is no shared technical definition, vendor terminology remains fluid, and public evidence does not justify a claim that the category is used uniformly across the industry.

Sources: s1, s2, s3

Limits and open questions

The word reasoning is a behavioral and product label, not proof that a model's intermediate process is faithful, complete, or human-like. The core sources report evaluations from the organizations that built the models, so broader independent testing remains important. Performance depends on the task, evaluation design, inference budget, and verification method. The available evidence supports examples from OpenAI and DeepSeek; it does not support saying that every major laboratory has adopted the same category or architecture.

Sources: s1, s2, s3

Related terms

References

Last updated: 2026-08-27

In the Skills Atlas

This term is also covered in the Skills Atlas as reasoning models skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as test time compute scaling skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as reinforcement learning from verifiable rewards skill.