acceptodds
Under review as a conference paper at ICLR 2027

With Tools or Because of Tools? Tracing Capability Gains in Multimodal Agents

Abstract

Multimodal agents increasingly rely on tools for image inspection and code execution, and their benchmark gains are often attributed to these interactions. Yet aggregate accuracy does not reveal whether tools expand the set of solvable problems or merely provide alternative paths to answers already within reach. We study DeepEyesV2, Thyme, and Qwen3-VL-8B-Instruct across 15 benchmarks, comparing each agent with its tool-free counterpart. We additionally train a tool-free multimodal reasoner from DeepEyesV2 data and edit tool-use trajectories to probe the roles of generated calls and returned results. Tool access yields no consistent performance advantage, while the trained reasoner matches or exceeds the two specialized agents on several benchmarks. Across DeepEyesV2 and Thyme, 93.2% and 96.0% of tool-enabled successes are reproduced by at least one non-tool reference. Some remaining successes persist after returned results are removed, whereas tool outputs judged novel, correct, and properly integrated often occur on problems already solvable without tools. The full tool-use loop yields clear gains in specific settings, but much of its apparent advantage remains attainable through non-tool reasoning, revealing a gap between tool-enabled success and tool-contributed capability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.