Reasoning Models
Models trained to reason before answering, emitting long internal chains of thought (OpenAI o-series, DeepSeek-R1, Claude extended thinking).
Test-time compute became a scaling axis, not just parameters.
LLMs post-trained to spend tokens thinking before answering, emitting long internal chains of thought checked by verifiable reward (OpenAI o-series, DeepSeek-R1, Claude extended thinking). The gains come from RL against outcome signals, not from bigger pretraining.
DeepSeek-R1 (Jan 2025) showed pure outcome-reward RL induces reasoning with no process supervision, and open weights collapsed the cost of reproducing o1-class results — reasoning stopped being a two-lab secret.
That the visible chain of thought is a faithful trace of the computation. It is a sampled sequence optimized for a correct final answer; models reach answers by paths the text doesn't show, so reading the CoT to audit safety or correctness is unreliable.
The frontier is reasoning models, so knowing when extra test-time compute pays for itself — and when it just burns tokens on easy queries — gets more valuable, not less.
Prerequisites
Reasoning models are transformer LLMs.
- mediumRLHF
Reasoning is elicited via RL post-training.
Recommended reference
Reviewed sources
Primary and first-party material reviewed for this editorial summary. These citations are separate from the AI consensus score above.
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs
Primary technical report on outcome-reward reinforcement learning and open reasoning models.
- Learning to reason with LLMs
First-party overview of the o1 reasoning-model approach and test-time reasoning behavior.
- Chain-of-Thought Monitorability: A New and Fragile Opportunity for AI Safety
Research on why visible reasoning traces are useful but cannot be treated as faithful explanations.