Latent Reasoning
Latent reasoning is multi-step inference performed through continuous internal representations instead of expressing every intermediate step as natural-language tokens. A model may feed a hidden state back as the next reasoning input, compress a textual trace into latent states, or mix latent and explicit steps. The term does not mean merely that neural networks have hidden activations; it denotes a designed mechanism that allocates intermediate computation in a non-textual representation space.
Origin and context
A November 2023 paper distilled explicit reasoning into hidden states without decoding intermediate steps. The December 2024 Coconut paper developed recurrent continuous thought, and a 2025 preprint survey organized several approaches under latent reasoning. CODI subsequently appeared at EMNLP 2025. Independent AAAI 2026 work introduced Dynamic Latent Reasoning with switching between discrete and continuous steps; ACL 2026 work studied interventions on continuous thought vectors. The category now spans multiple independently published mechanisms rather than naming Coconut alone.
Why it matters
Textual chains of thought consume tokens and force internal computation through a serial, human-readable channel. Continuous states can carry denser information and may reduce the number of decoded reasoning tokens. They also change observability: developers cannot inspect a vector sequence as easily as a written derivation. Latent reasoning therefore creates a tradeoff among inference cost, task performance, controllability, and auditability rather than a simple replacement for explicit reasoning.
Example
Consider a planning task with several plausible next moves. An explicit chain-of-thought model emits one textual step, commits it to context, and continues token by token. A Coconut-style model can pass a continuous thought state through another model step before decoding an answer; implicit-CoT and CODI-style systems instead learn hidden representations from explicit teacher traces. A model that silently uses ordinary transformer layers and then answers directly is not, by that fact alone, an implemented latent-reasoning system.
How it differs
Tool-Integrated Reasoning (TIR)
Tool-integrated reasoning interleaves model reasoning with observable calls to external executors or retrievers. Latent reasoning moves selected intermediate computation into continuous internal states. A system can combine both, but neither mechanism implies the other.
Test-time compute
Test-time compute is the broader practice of spending additional inference resources on a problem. Latent recurrence is one possible mechanism; sampling more textual answers or searching against a verifier can spend additional test-time compute without using latent states.
Maturity and evidence
Maturity is rated 3 for an established research category, not a production capability. Independent peer-reviewed work at EMNLP 2025, AAAI 2026 and ACL 2026 uses the same continuous-intermediate-computation meaning while studying different mechanisms. That sustained technical usage supports the lifecycle reassessment. The individual methods remain experimental: their task results do not establish broad deployment, comparable wall-clock savings or a generally superior replacement for explicit reasoning.
Limits and open questions
Fewer decoded tokens do not necessarily mean less total computation: latent iterations still execute model operations. The hidden representations are also harder to inspect than a textual derivation. ACL 2026's intervention study addresses that controllability problem in evaluated systems, rather than proving all continuous states are transparent. Comparisons must state the training procedure, model, tasks and reasoning budget; successes on particular benchmarks are not evidence of general reliability or production adoption.
Related terms
References
- Training Large Language Models to Reason in a Continuous Latent SpaceMeta, NYU, and UC San Diego / arXiv · 2024-12-09 · class A
- A Survey on Latent ReasoningIndependent multi-institution research team / arXiv · 2025-07-08 · class A
- CODI: Compressing Chain-of-Thought into Continuous Space via Self-DistillationAssociation for Computational Linguistics · 2025-11 · class A
- Implicit Chain of Thought Reasoning via Knowledge DistillationAllen Institute for AI, Microsoft, Johns Hopkins, and Harvard / arXiv · 2023-11-02 · class A
- ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem SolvingTsinghua University and Microsoft / arXiv · 2023-09-29 · class A
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model ParametersUniversity of California, Berkeley / arXiv · 2024-08-06 · class A
- Beyond Tokens: Dynamic Latent Reasoning via Semantic Residual RefinementTsinghua University and Kuaishou / AAAI · 2026-03-14 · class A
- Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to InterventionChang et al. / Association for Computational Linguistics · 2026-07 · class A
Last updated: 2026-09-05
This term is also covered in the Skills Atlas as reasoning models skill.
This term is also covered in the Skills Atlas as model training skill.