S24D: Generalizable Streaming Reconstruction for Sparse-View Dynamic Scenes
Abstract
High-quality 4D reconstruction from multi-view videos is fundamental to immersive volumetric video. However, existing dynamic novel-view synthesis methods often require densely sampled camera arrays and scene-specific optimization, making capture costly and introducing substantial processing latency. We propose SparseStream4D, a generalizable sparse-view streaming framework for low-latency dynamic 4D reconstruction. Its Sparse-view Anchor-driven Motion Network (SAM-Net) extracts multi-view 2D motion features and lifts them to compact 3D Gaussian anchors via visibility- and depth-aware aggregation. To preserve fidelity over long sequences, we introduce Interleaved Refinement Streaming, which periodically performs lightweight local refinement, and max-points bounded recalibration to control Gaussian growth and prevent sparse-view overfitting. A sparse-view flow consistency loss further enforces cross-view geometric coherence during training. Experiments on in-domain and cross-domain datasets demonstrate state-of-the-art rendering quality, 0.62-second per-frame reconstruction latency, and effective generalization to unseen scenes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.