With Tools or Because of Tools? Tracing Capability Gains in Multimodal Agents
Abstract
Multimodal agents increasingly rely on tools for image inspection and code execution, and their benchmark gains are often attributed to these interactions. Yet aggregate accuracy does not reveal whether tools expand the set of solvable problems or merely provide alternative paths to answers already within reach. We study DeepEyesV2, Thyme, and Qwen3-VL-8B-Instruct across 15 benchmarks, comparing each agent with its tool-free counterpart. We additionally train a tool-free multimodal reasoner from DeepEyesV2 data and edit tool-use trajectories to probe the roles of generated calls and returned results. Tool access yields no consistent performance advantage, while the trained reasoner matches or exceeds the two specialized agents on several benchmarks. Across DeepEyesV2 and Thyme, 93.2% and 96.0% of tool-enabled successes are reproduced by at least one non-tool reference. Some remaining successes persist after returned results are removed, whereas tool outputs judged novel, correct, and properly integrated often occur on problems already solvable without tools. The full tool-use loop yields clear gains in specific settings, but much of its apparent advantage remains attainable through non-tool reasoning, revealing a gap between tool-enabled success and tool-contributed capability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.