acceptodds
Under review as a conference paper at ICLR 2027

Query-Framed Spatial Reasoning via Spatial Latent Transformation in Video MLLMs

Abstract

Answering spatial questions from video requires using observations made at different positions and viewing directions. Remembering what was seen is not enough when a question asks for a relation from another viewpoint. Learning a scene representation is one way to support such reasoning. We investigate a direct alternative: adjust the observed visual evidence for answering, without requiring a video multimodal large language model (Video MLLM) to first generate a separate scene representation. We propose QUASAR (QUery-frAmed SpAtial Reasoning) framework, centered on Spatial Latent Transformation. Given estimated observation geometry and an explicit reference type and numerical pose, it adds bounded spatial corrections to the original visual features. The MLLM reads the full question and answers from this adjusted evidence; geometry remains required at inference. Spatial supervision, including four learned spatial summary tokens, accompanies supervised fine-tuning, followed by Group Relative Policy Optimization. QUASAR exceeds the original Qwen3-VL-8B baseline by 5.54 points on VSI-Bench and 19.88 points on VSTI-Bench; the historical VSI comparison uses different visual inputs and answer options. On both benchmarks, it also outperforms the Visual-only baseline (SFT+GRPO), which shares the backbone and data recipe but differs in pre-GRPO SFT exposure. Fixed-model interventions further support use of the visual corrections and, on parsable target-camera questions, the numerical reference. These results support direct visual-evidence adjustment as a viable route to spatial answering.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.