acceptodds
Under review as a conference paper at ICLR 2027

VisAtom: Do Multimodal Models Truly See Before They Think?

Abstract

Multimodal large language models (MLLMs) have achieved impressive scores on established reasoning benchmarks, yet recent evidence shows that many such benchmarks can be solved through textual shortcuts rather than genuine visual interpretation. This gap is critical because answer accuracy alone can overestimate capabilities and reward models that fail when visual conditions change. We introduce **VISATOM**, a methodology and benchmark for evaluating visual perception, reasoning with reduced textual cues, and generalization to modified problems. Starting from existing image-based problems, we extract *atomic facts*: elemental, directly observable statements describing content supported by an image, and use them to construct and check diagnostic tasks. These facts connect three evaluation tiers: **Perception Probing** tests basic visual observation, **Pruned Problems** remove image-redundant text, and **Edited Variants** programmatically alter problem instances to test generalization beyond public benchmark examples. Sourced from 2,711 challenging problems across 12 existing benchmarks spanning 8 domains, VISATOM provides a unified, multi-granularity diagnostic system on top of existing community benchmarks. Across ten frontier and open-weight MLLMs, VISATOM reveals persistent weaknesses in visual estimation (81–92%) and universal degradation after text pruning (4–11%). Edited variants further expose a 12.4% generalization gap for the strongest model (GPT-5.4), suggesting that strong reasoning performance does not always transfer to new problem variants.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.