acceptodds
Under review as a conference paper at ICLR 2027

Where Binding Begins: Atomic Feature Maps in Vision-Language Models

Abstract

Vision-language models (VLMs) describe complex scenes fluently, yet they fail at conjunctive search, where the target is defined by a combination of features, such as a particular color and shape, and each distractor shares one of them. This failure is a signature of the binding problem: keeping track of which features belong to which object. Humans show the same failure when they must respond before attention can select objects one at a time. Recent mechanistic interpretability work suggests that VLMs bind features through content-independent spatial retrieval mechanisms, but it remains unclear how the feature representations supplied to these mechanisms are themselves constructed. By intervening on the internal activations of Qwen2.5-VL-7B, we find that the searched-for features are gathered while the model reads the question (e.g., “Is there a green letter L?”) by attention heads at the color word and at the shape word that build maps of where each feature appears in the image. Patching these maps between displays makes the model attach a feature to the wrong object and report an absent target. These illusory conjunctions grow as objects get closer together, as in human crowding, and this proximity effect is carried by shape heads, not color heads. Supplying a clean shape map raises search accuracy from 32% to 94%, whereas a clean color map has little effect. Ablating the maps impairs counting but spares spatial-relation judgments, consistent with a separate indexing mechanism. These results suggest that, within a single forward pass, the model implements a parallel, feature-based stage of search but lacks the serial, object-by-object selection that reliable binding requires.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.