Knowing Without Deciding: The See–Decision Gap in Vision-Language Models
Abstract
Vision-language models (VLMs) can internally see the correct spatial relation yet still give the wrong answer. We term this the Encoding–Decision Gap: the ground-truth relation remains reliably decodable across many consecutive layers, but fails to determine the model's final response. Combining spatial readouts with causal interventions, we trace the gap to a dissociation between spatial encoding and decision control. Spatial information is broadly recoverable across intermediate states, yet causal influence over the answer is concentrated in a small subset of token–layer states, forming an encoding-to-decision bottleneck. Moreover, the directions that most strongly control the answer extend substantially beyond the relation-defined spatial subspace, while remaining relation-specific. We therefore use the model's own intermediate spatial estimate to select which relation to strengthen in the intermediate spatial states, and let the remaining computation propagate this change to the output. Without supplying an external answer or new spatial information, this intervention reconnects information already present in the model to a causally effective pathway, improving spatial reasoning across multiple datasets and model families. Our results show that spatial reasoning failures can stem not from missing representations, but from a failure to transfer encoded information into the decision, making encoding-to-decision transfer a concrete target for mechanistic analysis and intervention in VLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.