Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies
Abstract
Multimodal Large Language Models (MLLMs) now match or exceed human performance on many vision-language benchmarks. However, these evaluations typically rely on standard in-distribution data, leaving the robustness of MLLMs largely unexamined under prior conflict scenarios. To address this gap, we introduce VIA-Bench, a diagnostic benchmark that probes whether MLLMs remain grounded in objective physical reality under such conflicts. VIA-Bench covers six categories: color illusions, motion illusions, gestalt illusions, geometric and spatial illusions, general visual illusions, and visual anomalies. Through careful human-in-the-loop review, we construct over 1K high-quality question-answer pairs supporting both multiple-choice and open-ended evaluation. Extensive evaluation of 30 state-of-the-art MLLMs, including proprietary, open-source, and reasoning-enhanced models, uncovers significant vulnerabilities. Beyond aggregate accuracy, we further find that general-purpose Chain-of-Thought (CoT) prompting offers negligible robustness, often yielding “brittle mirages” where the model's logic collapses under illusory stimuli. Even task-specific CoT helps only marginally. To better characterize these failures, we introduce an error taxonomy that disentangles visual misperception, self-negation, and overthinking. These findings expose a structural vulnerability of current MLLMs under perceptual conflict and motivate future work on visually grounded reasoning. The benchmark data and code will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.