Correct After Rearrangement: A Tile-Permutation Test of Spatial Reliance in Vision-Language Benchmarks
Abstract
Vision-language models are scored on benchmarks such as MMStar, RealWorldQA, and ChartQA, where a score above a text-only baseline is often read as evidence that the model used the spatial arrangement of the image. Removing the image shows which scores need an image, not which properties of that image they need. A model could use the spatial arrangement of the scene, or collect evidence from separate regions independently of their layout. We separate these explanations by cutting each image into equal tiles, permuting the tiles at random, and passing the result to the model's own processor, which preserves the pixel content of every tile and disrupts inter-tile adjacency and tile location. We introduce persistence, the probability that a correct answer survives this intervention, and measure it for three open vision-language models on these three benchmarks. At a two-by-two grid, correctness is retained on about 65% to 80% of permutation trials among originally correct items, and on 44% to 79% at a four-by-four grid. This persistence exceeds content-removed baselines by 17 to 55 percentage points on a matched scale in all 27 model-benchmark-grid combinations, including every combination at the finest grid. Scrambling pixels inside the tiles while keeping the layout fixed reduces accuracy to near the text-only baseline, which places within-tile structure ahead of tile arrangement in this ordering. Most originally correct predictions therefore remain recoverable after the original coarse arrangement is disrupted. We release per-item records so that spatial reliance can be examined item by item.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.