acceptodds
Under review as a conference paper at ICLR 2027

Where Do Spatial Answers Come From? Tracing Utility Across the Visual Scene

Abstract

Vision-language models (VLMs) have made rapid progress on video-based spatial reasoning. However, it remains unclear which visual features these models rely on to generate their answers. Attention maps and representation probes reveal where a model attends and what information it encodes, but do not by themselves establish which feature directions a particular answer depends on. We introduce Utility Probe (U-probe), a training-free method that identifies an answer-conditioned feature subspace, termed Utility Space (U-space), by coupling visual representations with gradients of the answer. Deletion and steering interventions validate the causal role of U-space, whose contributions we trace across the visual scene. The leading feature directions in U-space are interpretable as spatial concepts, and subspaces for corresponding tasks agree across datasets. Deployed on VLMs finetuned on spatial tasks, U-probe reveals that scene-level spatial reasoning is supported by not only the queried objects but also features distributed over the scene, particularly on walls and floors, i.e., scene structure. During finetuning, regional utility after as few as 20 steps predicts later model performance, providing an early diagnostic of spatial SFT, whereas regional attention does not. Our results connect answer-conditioned feature subspaces to their use across the scene and identify scene structure as an important substrate for spatial reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.