Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction
Abstract
Streaming 3D reconstruction from extremely long videos requires online estimation of camera motion and scene geometry with bounded model memory and per-frame computation. Early streaming models use finite context buffers or compact recurrent states, yet their estimates often deteriorate as sequences grow. Recent methods improve long-horizon stability through persistent or multi-level long-range memory. We pursue a different route: local geometric prediction with reliable sequential composition. We present LC3R, a simple streaming model that caches KV features from only the preceding 11 frames. It predicts a point map in the current camera coordinate system and an adjacent-frame relative pose, keeping prediction targets local as the stream grows. Global poses and geometry are recovered by composing these local measurements. To ensure reliable long-horizon composition, a lightweight temporal rotation refiner improves relative rotations using recent visual and motion context, while a composition-aware pose loss directly supervises multi-step pose composition. On Oxford Spires, LC3R achieves an ATE of 4.35 m and an RPE-R of , reducing both errors by approximately 40% relative to the respective best prior streaming results. It runs at 24.45 FPS with 6.71 GB peak GPU memory, excluding input storage, on an NVIDIA H100 GPU.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.