Mamba-VGGT: Explicit Spatial State for Cross-Window Video Geometry
Abstract
Recovering camera motion and scene geometry from long videos requires strong multi-view reasoning and effective reuse of past observations. Pretrained geometric transformers provide powerful multi-view priors, but their global attention becomes increasingly costly as the input grows. Windowed inference limits this cost while interrupting feature-level communication across windows, which output-level geometric stitching cannot restore. We introduce Mamba-VGGT, a recurrent adapter that extends a frozen Visual Geometry Grounded Transformer with fixed-capacity cross-window memory. A Mamba-based two-dimensional selective scan (SS2D) jointly processes the incoming state and current patch tokens, producing coupled residual feature updates and memory writes. This design retains pretrained within-window attention while allowing subsequent windows to use historical context. State decay and periodic segment resets control state accumulation during long-sequence inference. On KITTI-360 sequences with 2,000 sampled frames, Mamba-VGGT reduces mean absolute trajectory error by 42.3% and mean LiDAR-supported reconstruction distance by 56.3% relative to LoGeR with sliding-window attention under the evaluated inference configuration. It also achieves lower trajectory error on VBR inputs of 8,000 and 10,000 frames under matched window and reset schedules. Mixer-replacement experiments favor SS2D over a parameter-matched pointwise MLP under the tested training setup, while state-propagation ablations show that carried context further improves trajectory estimation. Together, these results support joint spatial adaptation and cross-window memory as an effective approach to extending frozen geometric models for long-sequence camera estimation and reconstruction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.