Interleaved Thinking
Interleaved thinking is a model and runtime capability that allows reasoning steps between tool calls within one assistant turn. After receiving a tool result, the model can interpret the new evidence before deciding whether to call another tool, revise its plan or answer. It is more specific than making several calls in sequence: the defining feature is an intermediate reasoning opportunity informed by each result, subject to the model and API's thinking controls.
Origin and context
Anthropic announced extended thinking with tool use alongside Claude 4 on 22 May 2025, describing alternation between reasoning and tools. Moonshot AI independently described Kimi K2 Thinking, released on 6 November 2025, as interleaving chain-of-thought reasoning with function calls. A December 2025 independent arXiv preprint then used the same label in a tool-integrated research system.
Why it matters
Long tool workflows expose information that was unavailable when the first plan was formed: a search can return no evidence, code can fail, or an API can reveal a new constraint. Interleaving gives the model a structured chance to incorporate that observation before acting again. This can support error recovery and more selective tool use, but it also increases output-token use, context pressure and the number of consequential decision points. Product teams therefore need traces and evaluations that test the whole think-tool loop, not only the final answer.
Example
A research assistant searches for a paper, reads the returned abstract, notices that the result is a later survey, and changes its next query to the original title before drafting an answer. The reasoning between the read result and the second search is the interleaved step. If an orchestrator simply executes a fixed list of three calls without letting the model interpret intermediate results, it is sequential tool use but not interleaved thinking in this sense.
Sources: s4
How it differs
Tool-Integrated Reasoning (TIR)
Tool-integrated reasoning is the broader research and training paradigm in which tools participate in reasoning. Interleaved thinking describes the runtime arrangement that places reasoning between calls inside a turn. A TIR system may use that arrangement, but the two terms should not be treated as exact aliases.
Maturity and evidence
Maturity is rated 3. The capability has dated primary evidence, implementation in independently developed model families and use in an independent research preprint. It is no longer evidence from one launch. The rating remains below 4 because implementations differ across model families, evidence is concentrated in releases and one preprint, and comparative evidence for long production workflows is still limited.
Limits and open questions
Some providers expose thinking summaries rather than raw traces. Extra reasoning between calls can repeat mistakes, consume budget or trigger unnecessary actions. Vendor claims about hundreds of calls are model-specific and should not be generalized to the mechanism. Evaluations should report the model, tool interface, stopping policy and cost.
Related terms
References
- Introducing Claude 4Anthropic · 2025-05-22 · class A
- Kimi K2 Thinking model cardMoonshot AI · 2025-11-06 · class A
- MindWatcher: Toward Smarter Multimodal Tool-Integrated ReasoningIndependent researchers / arXiv · 2025-12-29 · class A
Last updated: 2026-09-04
This term is also covered in the Skills Atlas as reasoning models skill.
This term is also covered in the Skills Atlas as llm function calling skill.