AskGeo: Learning to Query Geometry Models for Spatial Reasoning
Abstract
Spatial understanding requires vision-language models (VLMs) to identify the geometric evidence relevant to a given question, across both static and dynamic scenes. Despite their strong semantic capabilities, existing VLMs often struggle to recover 3D structure and reason about spatial relations directly from visual observations. Meanwhile, pretrained feed-forward geometry models encode rich geometric knowledge, but existing methods transfer it by extracting geometric features independently of the question and only afterwards selecting or fusing them, so the geometry model performs the same computation regardless of what is asked. We observe that query-based geometry models such as D4RT already provide a retrieval mechanism: their decoder retrieves geometric information for arbitrary spatiotemporal queries. Building on this, we introduce AskGeo, a framework that lets a VLM learn to query a pretrained geometry model for the evidence a question requires. AskGeo projects question-conditioned VLM features into the native query space of the geometry decoder, which retrieves geometric evidence tailored to the question and returns it to the VLM as a few geometry tokens. Because the queries are formed from the VLM's own representation of the question, the retrieved tokens are aligned with how the VLM interprets the question, and what to ask is learned end to end from question-answer supervision. Experiments on the static VSI-Bench and dynamic DSR-Bench show that AskGeo achieves state-of-the-art performance on both benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.