Look Once: A mid-depth commitment window explains Spatial Instability in VLMs
Abstract
Vision–language models (VLMs) perform inconsistently in spatial reasoning where some lateral relations are near ceiling, while real-world height and orientation remain at or below chance. Their predictions also change under geometric transformations that preserve the correct response. What remains unclear is where the decision forms and whether richer visual inputs, spatial post-training, and explicit reasoning alter this process or improve its reliability. We combine residual-state patching across token regions and depth with an answer-preserving scene reversal that jointly reverses the spatial relation and option order, changing the scene while keeping the correct answer slot fixed. This separates scene-dependent information transfer from an already resolved verdict. Across three architectures, we identify a narrow commitment window where option-token states become verdict-like as causal influence from the vision block declines. Restoring early visual states produces substantial scene-specific transfer before this boundary but little afterward, supporting reduced downstream use rather than content loss alone. Moreover, 88-98% of the most influential attention heads and MLPs lie downstream of commitment. Spatial post-training preserves the broad relay organization, while explicit reasoning improves accuracy by gaining 4.1%, yet prediction flips under answer-preserving reflection increase from 12.3% to 15.0%. These findings shift the design target from supplying more visual evidence to controlling how that becomes a verdict, the commitment window marks where spatial decisions take shape, while most influential downstream computation operates on the verdict already formed.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.