acceptodds
Under review as a conference paper at ICLR 2027

Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies

Abstract

Multimodal Large Language Models (MLLMs) now match or exceed human performance on many vision-language benchmarks. However, these evaluations typically rely on standard in-distribution data, leaving the robustness of MLLMs largely unexamined under prior conflict scenarios. To address this gap, we introduce VIA-Bench, a diagnostic benchmark that probes whether MLLMs remain grounded in objective physical reality under such conflicts. VIA-Bench covers six categories: color illusions, motion illusions, gestalt illusions, geometric and spatial illusions, general visual illusions, and visual anomalies. Through careful human-in-the-loop review, we construct over 1K high-quality question-answer pairs supporting both multiple-choice and open-ended evaluation. Extensive evaluation of 30 state-of-the-art MLLMs, including proprietary, open-source, and reasoning-enhanced models, uncovers significant vulnerabilities. Beyond aggregate accuracy, we further find that general-purpose Chain-of-Thought (CoT) prompting offers negligible robustness, often yielding “brittle mirages” where the model's logic collapses under illusory stimuli. Even task-specific CoT helps only marginally. To better characterize these failures, we introduce an error taxonomy that disentangles visual misperception, self-negation, and overthinking. These findings expose a structural vulnerability of current MLLMs under perceptual conflict and motivate future work on visually grounded reasoning. The benchmark data and code will be released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.