Glossary · term

Test-time compute

Test-time compute is the computation allocated after a prompt arrives and before an answer is finalized. For a language model, scaling it can mean generating several candidate solutions and selecting among them with a verifier, or allowing a trained model to use a longer adaptive reasoning process. It is a resource-allocation strategy, not a model family or a training algorithm: the inference budget can vary while the underlying model remains the same.

Training2024-08-06Wave 1 · 2023Maturity: 3/5

Origin and context

In August 2024, Snell and colleagues studied how additional inference computation should be allocated for difficult LLM prompts. They compared verifier-guided search with an approach that adaptively changes the model's response distribution, and argued that the effective strategy depends on problem difficulty. In September 2024, OpenAI explicitly separated train-time reinforcement learning from time spent thinking at test time when reporting the behavior of o1. These sources document a research formulation and a deployed example; they do not establish that either group coined the general phrase.

Sources: s1, s2

Why it matters

Test-time compute moves part of the capability-and-cost decision from model training to each inference request. A system can reserve a larger budget for a hard mathematics, coding, or planning problem without paying that cost on every simple query. Snell et al. found that compute allocation should be adapted to the prompt and reported conditions in which a smaller model with additional inference work surpassed a much larger model at matched FLOPs. OpenAI separately reported that o1 performance improved with more time spent thinking. These results make latency, cost, verification quality, and task difficulty joint design variables rather than afterthoughts.

Sources: s1, s2

Example

For a difficult contest-math question, a test-time scaling system might generate multiple proposed proofs, score intermediate steps with a process-based verifier, and spend the remaining budget refining the strongest path. Another system may allocate a longer internal reasoning interval to the same prompt. For a routine formatting request, both strategies may add cost and delay without a meaningful benefit. The useful decision is therefore not simply whether to think longer, but how much computation to allocate and which search or reasoning mechanism is appropriate for this particular input.

Sources: s1, s2

How it differs

Reasoning models

A reasoning model is a category of model trained and presented for multi-step problem solving. Test-time compute is the inference resource or procedure applied to a request. Reasoning models often expose a controllable thinking budget, but test-time scaling can also search or rerank outputs from models not marketed as reasoning models.

Reinforcement Learning with Verifiable Rewards (RLVR)

RLVR is a training-time method that uses automatically checkable rewards on tasks such as mathematics or code. It can help produce reasoning behavior, as the DeepSeek-R1 work illustrates, but it occurs before deployment. Test-time compute concerns what happens after the trained model receives a prompt.

Maturity and evidence

Maturity is rated 3 rather than 4. The concept has a clear academic formulation and an independently documented production-model example, and related work appears in more than one organization. However, the best allocation method remains task-dependent, terminology overlaps with inference-time scaling, and the evidence does not support treating increased inference compute as a universal improvement.

Sources: s1, s2, s3

Limits and open questions

More test-time compute does not guarantee a better answer. Snell et al. report that strategy effectiveness varies with prompt difficulty and the base model's initial chance of success. Longer reasoning also increases latency and operating cost, while verifier-guided search depends on verifier quality. Vendor evaluations can demonstrate a system under stated settings, but they do not establish the optimal budget for other models, tasks, or deployment constraints.

Sources: s1, s2

Related terms

References

Last updated: 2026-08-27

In the Skills Atlas

This term is also covered in the Skills Atlas as test time compute scaling skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as reasoning models skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as reinforcement learning from verifiable rewards skill.