Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning
Abstract
Video spatial reasoning requires accumulating viewpoint-dependent evidence over time while retaining information useful to the question being asked. Existing spatial video-language models improve geometric perception and long-range context modeling, but often treat memory as a generic temporal cache, which can introduce redundant or irrelevant evidence and weaken long-horizon reasoning. We propose Q-GeoMem, a question-guided geometric memory framework for video spatial reasoning. Q-GeoMem injects camera-conditioned geometry into visual tokens and maintains a Fine-Grained Context Bank for recent dense context and a Semantic-Geometric Evidence Bank for compact long-range evidence. Each candidate frame is scored by calibrated question relevance and novelty with respect to the active bank, and the resulting utility drives both memory replacement and reading. Experiments on two in-domain and five out-of-distribution benchmarks show that Q-GeoMem achieves state-of-the-art performance in the evaluated settings, and controlled analyses validate question-guided geometric evidence selection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.