MAVER: Memory-Augmented View-Decoupled Evidence Routing for Video Spatial Reasoning
Abstract
Video spatial reasoning requires integrating spatial cues that are distributed across time and viewpoints. Recent geometry-grounded video-language models provide increasingly rich reconstruction and camera-aware features, yet these dense representations are commonly compressed before language reasoning, which can obscure question-relevant spatial evidence. We introduce MAVER, a structured evidence interface that operates on high-resolution geometry-grounded video features before spatial pooling. MAVER first conditions current spatial features on previously observed context, then organizes them into cross-view shared and observation-specific evidence, and finally performs question-conditioned sparse routing. The selected evidence follows a dual-path injection scheme that both guides the dense visual stream before pooling and preserves selected evidence through a direct projected bypass to the language model. On VSI-Bench and VSTI-Bench, MAVER achieves average scores of 64.4 and 61.6, respectively, including a 7.1-point improvement over the controlled geometry-grounded baseline on VSI-Bench. These results suggest that structured evidence transformation before language reasoning provides an effective way to improve geometry-grounded video spatial reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.