acceptodds
Under review as a conference paper at ICLR 2027

A Mechanistic View of Spatial Reasoning in Vision-Language Models

Abstract

Spatial reasoning is an emerging capability of Vision Language Models (VLMs). As VLMs see broader deployment in real-world settings requiring reliability and interpretability, understanding spatial reasoning mechanisms is critical. Through a mechanistic view, we find evidence that spatial reasoning emerges from a coordinated division of labor in VLMs defined by vision encoder (VE) components that produce tokens consumed by large language model (LLM) components. Specifically, we show that spatial reasoning in VLMs depends strongly on positional encodings in the VE and on interactions between background tokens, while the LLM selectively integrates this information from background image tokens and foreground image tokens through a sparse set of functionally specialized attention heads. Despite our probing task’s focus on queried objects, we find that background tokens also play a critical role by providing spatial context that is actively leveraged by the model during spatial reasoning. Based on these findings, we show that mechanistic interpretability can guide visual token compression. The performance on real-world datasets and diverse spatial tasks suggests that the discovered pathways extend beyond the original analysis setting, enabling more efficient visual token processing without sacrificing reasoning performance. Overall, our results improve understanding of the mechanisms underlying spatial reasoning in VLMs, and we demonstrate the resulting insights for inference efficiency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.