acceptodds
Under review as a conference paper at ICLR 2027

Remember, Revisit, Revise: Scaffold Memory for Long-Horizon 3D Geometry

Abstract

In 3D computer vision, geometry foundation models estimate camera poses and 3D scene structure from video sequences, but maintaining consistent reconstructions over long horizons remains challenging. Feed-forward models process long sequences as independent, fixed-length frame windows and subsequently align their predictions, causing pose and reconstruction errors to accumulate across windows. Streaming models retain information from previous frames to inform future predictions, but cannot revise earlier estimates when subsequent observations reveal errors. To address these limitations, we introduce Scaffold Memory for Revision (SMoRe), a framework that equips existing geometry foundation models with a revisable external memory for long-horizon inference. Inspired by hippocampal associative memory in biological brains, which links spatial context with sensory experience, SMoRe associates where an observer is with what it observes, enabling past observations to be stored, retrieved, verified against new evidence, and revised when inconsistencies are detected. Crucially, SMoRe integrates with existing geometry foundation models while keeping them frozen, requiring no additional training or fine-tuning. We evaluate SMoRe with eight competitive geometry models on seven datasets spanning camera pose estimation, multi-view depth estimation, 3D point-cloud reconstruction, simultaneous localization and mapping (SLAM), and novel view synthesis. Across all the evaluated geometry models, SMoRe consistently improves long-sequence camera pose estimation, reduces multi-view depth error by up to 51%, reduces Chamfer distance for 3D point-cloud reconstruction by up to 44%, and achieves average trajectory error as low as 2.9 cm in long-horizon SLAM. By retaining corrected camera poses and observations, SMoRe further improves the fidelity and spatial consistency of large-scale 3D scene reconstruction. These results establish revisable external memory as a flexible interface for extending pretrained geometry models beyond fixed-context prediction toward iterative, error-correcting inference over long visual sequences. Code, data, and models will be publicly available upon publication.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.