acceptodds
Under review as a conference paper at ICLR 2027

VTC-Bench: Evaluating Agentic Multimodal Models via Compositional Visual Tool Chaining

Abstract

Recent advancements extend Multimodal Large Language Models (MLLMs) beyond passive visual perception to actively composing external tools for complex visual tasks. Evaluating these capabilities requires examining how models select visual operations, compose them into multi-step workflows, and interpret intermediate outputs. We introduce VisualToolChain-Bench (VTC-Bench), a benchmark for evaluating multi-step visual tool composition in MLLMs. Our framework features 35 diverse visual operations and 680 curated problems structured across a nine-category functional taxonomy, each paired with reference execution trajectories that serve as diagnostic anchors for fine-grained analysis. Crucially, VTC-Bench supports a dual-paradigm evaluation protocol: a code-driven mode that tests programmatic tool orchestration, and an interface-driven mode that assesses structured tool invocation through predefined tool-sets. This design enables systematic diagnosis of where model capabilities diverge across execution paradigms. Extensive experiments on 19 leading MLLMs reveal critical limitations: models struggle with compositional visual tool use, with the best model achieving 51% accuracy. Multi-tool composition remains a fundamental bottleneck. Models concentrate their calls on a small subset of operations and make limited use of reference tools even when explicitly provided. By identifying these challenges, VTC-Bench provides a rigorous foundation for developing more capable visual agentic models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.