CoSTA4D: Consistent Spatio-Temporal Aggregation for Streaming 4D Human Reconstruction from Uncalibrated Sparse-View Videos
Abstract
Streaming human reconstruction from sparse, uncalibrated multi-view videos requires both geometric accuracy and temporal consistency. While human motion and visibility vary over time, the capture cameras remain fixed; however, changing visual evidence can still cause unstable camera and depth predictions, leading to geometric jitter and rendering flicker. This suggests that historical observations should not only provide temporal context for current reconstruction, but also accumulate evidence about the invariant capture geometry. We present CoSTA4D, a streaming feed-forward framework built on this principle. CoSTA4D combines anchor-context temporal aggregation with reliability-aware camera refinement, enabling historical context to stabilize geometry inference while repeated camera observations progressively improve the shared camera estimates and back-projected geometry. Correspondence-guided Gaussian propagation further encourages continuity of the dynamic representation across frames. The framework uses only current and past observations without per-sequence optimization. Experiments on DNA-Rendering and self-captured videos demonstrate state-of-the-art reconstruction quality and improved temporal consistency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.