Not Every Tool Call Helps: Fine-Grained Credit Assignment for Efficient Tool-Integrated Reasoning
Abstract
Tool-Integrated Reasoning (TIR) enables large language models (LLMs) to solve complex reasoning problems with external tools, yet existing reinforcement learning methods largely rely on coarse-grained trajectory-level rewards and overlook the utility of individual tool calls. We systematically study tool-use behavior and identify pervasive tool overuse: models often invoke tools even when their current reasoning states are already sufficient for correct solutions. To quantify this behavior, we introduce an early-stopping estimator that measures the observable correctness gain of each tool call by comparing outcomes immediately before and after invocation. Our analysis reveals that many tool calls are redundant or harmful, driven by persistent tool-calling bias and limited adaptation to prior tool utility. Based on these findings, we propose GAPO (Gain-Aware Policy Optimization), which converts tool-call gains into fine-grained local advantages and combines them with trajectory-level supervision to optimize tool-calling behavior. Across mathematical reasoning and code-generation benchmarks with Qwen3 models from 4B to 32B, GAPO consistently improves reasoning performance and tool-use efficiency. On Qwen3-8B, it improves average mathematical reasoning accuracy by 13.77 percentage points while reducing tool calls by 80.63%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.