Glossary · term
Tool-Integrated Reinforcement Learning / TIR-RL
RL for reasoning models in which thinking steps are coupled with tool calls (code, computation, search), moving Tool-Integrated Reasoning from prompting into post-training. The SimpleTIR paper stabilizes multi-step training by removing trajectories with void turns (steps containing neither code nor an answer) from the policy update while keeping them in the advantage estimation; starting from a Qwen2.5-7B base it reaches 50.5 on AIME24. Xue, Zheng, Liu, et al., ICLR 2026.
Training2026Wave 3 · 2025–26Maturity: 1/5
Maturity rationale
speculative / early neologism (warning)
References
Author: Społeczność / Anonimowi