acceptodds
Under review as a conference paper at ICLR 2027

Same same, but different: VLMs Do Not Recognize Features Equally Anywhere in an Image

Abstract

Vision-Language Models (VLMs) have achieved impressive performance on a wide range of tasks requiring visual understanding. Despite this progress, their predictions remain surprisingly brittle and sensitive to small spatial variations. Using synthetic experiments that isolate specific visual factors, we systematically study anisotropic feature recognition in VLMs. We find that modern VLMs reliably recognize simple atomic features contained within single patches, but continue to struggle with more complex atomic and compositional features. Crucially, the same feature can be recognized with considerably different fidelity depending on its position in the input image, with spatial anisotropy patterns that vary across feature types and models. We further show that fragmentation across patches is rarely harmful, while upsampling a feature over more patches substantially improves recognition, particularly for compositional features. Conversely, visual distractors and longer vision-token sequences degrade feature recognition and introduce additional positional biases. Taken together, our results reveal that feature recognition in VLMs remains strongly anisotropic and sensitive to the allocation of visual information across tokens.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.