MV3R: Long-Sequence Multi-View 3D Reconstruction
Abstract
Reconstructing 3D scenes from long-horizon multi-camera streams requires processing thousands of images without sacrificing global geometric consistency. Existing feed-forward methods, however, typically process all input views in a single forward pass and therefore scale poorly to long sequences, while streaming methods rely on a bounded state that may discard fine-grained cross-view evidence and accumulate drift over time. To overcome this trade-off, we introduce MV3R, a scalable window-based framework that processes a sequence as overlapping local windows and propagates context through their shared observations, thereby preserving continuity without jointly processing the entire sequence. Building on this windowed design, MV3R incorporates patch-level Plücker ray encodings to expose the spatial configuration of the fixed camera array and improve reconstruction within each window. The shared observations between adjacent windows then serve a second purpose by establishing deterministic correspondences for connecting local reconstructions. To robustly exploit these correspondences, a fixed-capacity scene memory predicts their reliability and guides a two-stage Sim(3) registration that performs confidence-weighted coarse alignment followed by residual refinement. Together, these designs enable MV3R to scale to long multi-camera sequences with bounded memory while bringing all local reconstructions into a geometrically consistent global coordinate system. To address the lack of benchmarks for indoor long-sequence reconstruction with fixed multi-camera arrays, we further introduce MV3D for evaluating geometric reconstruction quality and long-term trajectory stability. Extensive experiments show that MV3R significantly outperforms existing methods in reconstruction and trajectory estimation across four indoor datasets. Specifically, relative to the strongest competing result for each reconstruction metric on each test subset, MV3R reduces point-cloud Accuracy and Completeness errors by 22.5% and 22.1% on average, respectively. For trajectory estimation, MV3R reduces average ATE from 0.204 for the strongest baseline to 0.095, corresponding to an approximately 53.4% reduction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.