Learn How Before When: Decoupling Capability and Routing for Visual Tool Use
Abstract
Tool-augmented reasoning requires learning to answer in both the direct and tool modes and to route each question appropriately. Existing reinforcement learning with verifiable rewards (RLVR) methods often train mode capability and routing jointly and can exhibit phenomena like tool overuse or tool collapse, yet the connection between these behaviors and joint training remains unclear. We analyze GRPO advantages and their effects on one-step changes in mode capability and routing. The analysis shows how a mode can lose sampling probability while its capability is still improving, and why reward incentives can favor tool use without a corresponding accuracy benefit. Training diagnostics under representative reward settings further illustrate sharp declines in mode frequency and persistent tool calls under sustained balancing. Motivated by these findings, we propose a two-stage framework that separates mode capability learning from routing. We first use GRPO to train a shared expert under explicit mode instructions, then compare its performance across modes to guide an autonomous student through on-policy self-distillation. This design transfers mode capabilities and question-specific routing without additional routing rewards or large-scale supervised tool-use trajectories. Experiments show accuracy gains of up to 14.13 percentage points over Qwen2.5-VL-7B on high-resolution perception benchmarks. Distillation process further substantially reducing tool use rate, supporting more effective tool-use decisions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.