acceptodds
Under review as a conference paper at ICLR 2027

An Exam for Active Observers

Abstract

Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. Whether today's multimodal large language models (MLLMs) exercise active observation is an empirical question that current vision-language benchmarks do not answer. We introduce ActiveVision, a benchmark that makes active observation measurable for MLLMs, comprising 17 tasks across 3 categories. Tasks are designed to force repeated visual perception rather than a single static description. Frontier MLLMs collapse on ActiveVision: the highest-scoring model we evaluate, GPT-5.6 Sol, solves at most 23.5% of items at any reasoning-effort tier, and even Claude Fable 5, despite its strong results on reasoning and coding leaderboards, solves just 5.9% at its highest tier, far behind three human participants who average 96.1%. Furthermore, much of the gap persists even when models write and run their own vision code. Such code is unreliable on realistic imagery, and catching its failures itself requires the active perception the models lack. Together, these results indicate that current MLLMs lack robust active visual observation, motivating architectures and training objectives that close the perception–reasoning loop.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.