acceptodds
Under review as a conference paper at ICLR 2027

Reported Safe, Blind in Practice: A Cross-Modal Conflict Benchmark for the Action Layer of Multimodal Agents

Abstract

Multimodal agents increasingly act on values they read from images, so the critical failure is not a wrong caption but a wrong action taken when a textual instruction contradicts the visual evidence. Prior cross-modal conflict studies operate at the answer layer and report modality-preference or "safety" rates, which cannot tell a grounded correction apart from a mere refusal. We introduce an action-layer conflict benchmark of 1,500 programmatically constructed items from 750 TextVQA and ChartQA image-question pairs, split evenly into conflict and matched agreement conditions, and a grounding-certifying metric, Visual Accuracy, that credits an agent only when it overrides the mistaken instruction and reports the correct visible value. Evaluating six state-of-the-art models, we find that Visual Accuracy collapses to at most 0.40 even though blind-follow rates are near zero, that this collapse and its model ordering are preserved across both distributions, and that the high "safety" of standard metrics is overwhelmingly refusal rather than perception. We then build a verification ladder and show that an inference-time locate-read-compare-decide chain significantly improves Visual Accuracy for every model (GPT-5.5: 0.14 to 0.84) and rescues even the hardest model, while we document its cost of trading hollow safety-by-refusal for action. Our benchmark provides a faithful yardstick for multimodal agents that must act on what they see.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.