Glossary · term

Tool-Integrated Reasoning (TIR)

Tool-integrated reasoning, or TIR, is a reasoning pattern in which a language model interleaves natural-language deliberation with calls to external tools and then incorporates returned results into subsequent reasoning. The tools may calculate, execute code, retrieve evidence, or verify formal steps. TIR describes the trajectory format and capability, not a particular training algorithm: imitation learning, reinforcement learning, preference optimization, or prompting can each be used to elicit it.

Agents2023-09-29Wave 3 · 2025–26Maturity: 3/5

Origin and context

The September 2023 ToRA preprint named a Tool-integrated Reasoning format for mathematical problem solving, contrasting it with rationale-only and program-only approaches. ToRA trained models on interactive trajectories that alternated rationales and program execution and was later published at ICLR 2024. Independent work subsequently used TIR with Lean for partial proof validation, while a 2025 empirical study treated it as a broader paradigm and compared tool-enabled models with text-only counterparts across several reasoning categories.

Sources: s1, s2, s3, s4

Why it matters

A model's parametric knowledge and arithmetic are imperfect, while external systems can supply current evidence or exact execution. TIR lets the model decide when an outside operation is useful, translate a subproblem into a tool request, and reason over the observation rather than merely append a tool result. This creates capabilities that neither fluent text generation nor one-shot program synthesis provides alone, but it also makes correctness depend on orchestration, tool availability, and the model's interpretation of outputs.

Sources: s1, s2, s3, s4

Example

For a geometry problem, a TIR trajectory can first derive a relationship in words, call a symbolic solver to simplify an expression, inspect the returned value, and revise the proof. For a research question, the analogous pattern might formulate a search, retrieve evidence, and continue reasoning with citations. Merely exposing a calculator function is not sufficient: if the model never integrates the observation into a multi-step trajectory, the system supports tool use but has not demonstrated tool-integrated reasoning.

Sources: s1, s2, s3, s4

How it differs

Tool Use and Function Calling

Function calling is an interface for producing structured tool requests. TIR is the higher-level reasoning pattern that decides, sequences, and learns from calls. A single API call can be function calling without a tool-integrated reasoning trajectory.

Latent Reasoning

Latent reasoning performs selected intermediate computation in continuous internal states. TIR brings observations from external systems into the reasoning trace. They are complementary but independently defined mechanisms.

Maturity and evidence

Maturity is rated 3. TIR has a peer-reviewed foundational system, independent peer-reviewed reuse, and later cross-domain analysis. It remains a research paradigm rather than a standardized runtime contract; tool sets, trajectory formats, training objectives, cost measures, and benchmarks vary substantially.

Sources: s1, s2, s3, s4

Limits and open questions

The cited studies report execution errors, reasoning errors, unnecessary or failed tool use, and sensitivity to tool-call budgets. A model can construct a bad request, misread a returned result, or gain from a strong executor without improving its unaided reasoning. Evaluation should therefore report tool availability, call budget, failure rates, and a comparable no-tool baseline.

Sources: s1, s3, s4

Related terms

References

Last updated: 2026-09-04

In the Skills Atlas

This term is also covered in the Skills Atlas as reasoning models skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as code execution agents skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as agentic planning task decomposition skill.