Fast and stable long-context 3D reconstruction using Recon-Aware Memory Bank
Abstract
Feed-forward 3D reconstruction models such as VGGT provide high-quality geometry and camera estimation without per-scene optimization, but scaling them to long visual sequences remains challenging. Dense models incur quadratic attention computation as context grows, while streaming variants still accumulate increasingly large historical context. We argue that efficient context reduction should be both reconstruction-aware and architecture-decoupled: selected frames should preserve downstream reconstruction quality without relying on backbone-specific cache or attention structures. To this end, we introduce Reconstruction-Aware Memory Bank (RAMB), a plug-in framework that uses frame-associated scene tokens as a common representation across VGGT-style backbones. RAMB combines a lightweight Rank Model trained with downstream reconstruction outcomes to identify high-utility frames with scene-token similarity to retain complementary observations and remove redundant context. We evaluate RAMB on VGGT- and StreamVGGT across ScanNet, ETH3D, and KITTI-VO. The selected frames are highly informative: reconstructing from only about 3% of a 400-frame sequence, VGGT- matches or exceeds full-sequence reconstruction accuracy (ACC). For chunked inference, RAMB carries informative context across chunks, enabling 540-frame reconstruction on a single 24 GB GPU, where full-sequence inference requires over 60 GB, with comparable quality. For StreamVGGT, RAMB bounds the historical context to 20.9 of 240 frames on average while maintaining comparable quality, and reduces forward latency and peak memory by .
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.