acceptodds
Under review as a conference paper at ICLR 2027

Test-time reasoning with computational evidence for spatial question answering

Abstract

Vision-language models (VLMs) have made substantial progress, while the growing interest in embodied AI underscores the need for reliable spatial reasoning. Existing approaches have attempted to improve spatial reasoning by enhancing visual reasoning capabilities or augmenting VLMs with external visual tools. However, visual reasoning can be unreliable when the VLM lacks sufficient spatial sensing capability. Although tool augmentation can potentially address this limitation, invoking a large set of tools or reconstructing an entire 3D scene is often computationally expensive, inefficient, and unnecessary. Our analysis reveals a capability imbalance: VLMs show strong object-level perception and reasoning but struggle with spatial sensing, suggesting that explicit geometric evidence can help bridge this gap. We propose Test-Time Spatial Reasoning (TTSR), a training-free framework that supports spatial question answering through computational evidence. TTSR constructs a question-specific geometric plan, then grounds the entities and acquires the spatial information using the VLM and minimal auxiliary tools. It performs reference-frame transformations and compositional spatial computations, producing explicit evidence that the VLM interprets using its own reasoning capabilities to get the answer. Experiments on recently proposed challenging spatial reasoning benchmarks demonstrate that TTSR substantially outperforms existing state-of-the-art methods and generalizes consistently across diverse spatial reasoning settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.