FLRT: Learning Efficient Tool Use from Tool-Sufficient States
Abstract
Tool-augmented reasoning substantially extends the capability of large language models to solve complex problems, but frequent tool use also incurs additional interaction and execution costs. We observe that tool necessity changes dynamically along a reasoning trajectory: a model may rely on external tools to acquire critical information early on, yet become capable of completing the task independently at later states. To characterize this transition, we introduce Fork Exploration, which disables future tool access from real intermediate states and observes whether the model can still reach a correct solution, yielding a state-level signal of tool sufficiency. Building on this signal, we propose FLRT, a two-stage training framework that learns more efficient tool-use policies through fork-guided supervision and fork-aware reinforcement learning. Experiments across three mathematical reasoning benchmarks show that FLRT improves average accuracy over the base model while reducing tool calls by about one-third. Further analysis shows that FLRT primarily reduces tool interactions after the model has already acquired sufficient information, suggesting that explicitly modeling within-trajectory changes in tool necessity is an effective approach to improving the efficiency of tool-augmented reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.