acceptodds
Under review as a conference paper at ICLR 2027

M³AgentBench: Benchmarking Multi-Turn Agent–User Interaction in Multimodal Tasks

Abstract

Real-world agents must continue working as users intervene after the initial request, yet current multi-turn evaluations often focus on specific interaction phenomena rather than systematically diagnosing how agents respond to different forms of intervention. We introduce M³AgentBench, a benchmark for Multimodal, Multi-Turn, and Multi-Intervention agent evaluation. It contains 60 human-curated multimodal workflows across six domains and covers five representative user interventions: Clarification, Incremental, Revision, Topic Switching, and Adversarial Pressure. Each task requires persistent tool use and produces verifiable artifacts, while state-constrained user simulation and process- and result-aware grading enable reproducible evaluation of both interaction behavior and final outcomes. We evaluate frontier agents and further perform interaction-conditioned failure analysis to diagnose how agents fail beyond aggregate scores. Our study is designed to reveal whether similar overall performance can conceal distinct failure mechanisms and whether correct final artifacts can coexist with flawed interaction processes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.