Agentic-MME : What Agentic Capability Really Brings to Multimodal Intelligence?
Abstract
Multimodal Large Language Models (MLLMs) are rapidly evolving from passive observers into active agents. These agents increasingly solve problems through Visual Expansion, actively invoking visual tools to transform images. Furthermore, as real-world tasks often require information beyond visual content, modern systems combine these visual operations with Knowledge Expansion via open-web search. However, existing evaluations fall short in three critical aspects. (i) Lack flexibility and comprehensiveness in tool integration that supports heterogeneous tool interfaces. (ii) Test image tool use and web search separately, leaving their synergy unexplored. (iii) Evaluate primarily by the correctness of the final answer. Consequently, they cannot verify whether the tools were actually invoked, applied correctly, or used efficiently. To answer what agentic capability truly brings to multimodal intelligence, we introduce Agentic-MME, a process-verified benchmark for Multimodal Agentic Capabilities. Agentic-MME contains 418 real-world tasks across 6 domains and 3 difficulty levels designed to evaluate capability synergy, featuring over 2,000 stepwise checkpoints that average more than 10 person-hours of manual annotation per task. Each task is paired with (i) a unified evaluation framework that supports both sandboxed code execution and structured tool APIs, and (ii) a human reference trajectory annotated with stepwise checkpoints along dual-axis: S-axis and V-axis. To enable true process-level verification, we move beyond final answers by auditing fine-grained intermediate states. We also quantify efficiency via an overthinking metric relative to human trajectories. Experimental results show that the best model Gemini3-pro achieves an overall accuracy of 56.3% and this score falls significantly to 23.0% on Level-3 tasks, underscoring the difficulty of real-world multimodal agentic problem solving.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.