Solvable, Yet Still Wrong: An Axis-Wise Perspective on Spatial Reasoning in Vision-Language Models
Abstract
Despite significant progress in Vision-Language Models (VLMs), spatial reasoning, which involves inferring the spatial relationships between objects, remains challenging. Existing benchmarks typically report a single aggregated accuracy per task, making it difficult to determine whether a spatial relation is solvable from the available geometric evidence or how models fail. In practice, different spatial axes (e.g., horizontal, vertical, and depth) require different geometric evidence. We therefore introduce an axis-wise analysis that examines, for each relation, what information makes it solvable and how models fail when relevant information is supplied. Our analysis reveals a gap between what VLMs can spatially infer and what they ultimately answer: models can possess the relevant geometric information yet select a relation from the wrong axis. Through controlled analysis on real-world datasets, we reveal three key findings: i) Horizontal, vertical, and depth relations differ substantially in their geometric solvability and in how accurately VLMs can solve them from the available evidence, showing that spatial reasoning is not uniform across axes. ii) We observe that, when VLMs make mistakes, they often select a relation from an incorrect axis, even when at least one geometrically valid spatial relation is among the options. Interestingly, restricting the multi-axis options to the correct axis substantially improves accuracy, highlighting axis selection as an important error source. iii) This axis-wise failure can be alleviated by supplying axis-specific geometric evidence. Providing geometry as the relation implied by grounded objects significantly improves accuracy, whereas directly providing object coordinates sometimes even hurts performance, indicating that how geometric evidence is represented matters. Based on these findings, we introduce AxisCue, a training-free, axis-aware inference method that allows axis-conditioned spatial reasoning. AxisCue identifies the queried axis, grounds the reference objects, and converts them into axis-specific geometric evidence, eliciting relevant geometric information from the model to improve accuracy. Overall, our analysis demonstrates that an axis-wise perspective exposes failure modes hidden by aggregated spatial accuracy and provides a principled basis for deriving simple, targeted improvements to spatial reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.