ReStream: Fast Streaming Reconstruction with Incremental Dynamic Token Updates
Abstract
Volumetric video reconstruction aims to recover dynamic 3D scenes from synchronized multi-view videos, which are commonly captured by fixed camera rigs. Recent Large Reconstruction Models (LRMs) offer a promising feed-forward solution, yet applying them frame by frame incurs substantial redundant computation, as large portions of the scene remain unchanged throughout the video. We introduce \ourname, a training-free framework that converts frame-wise LRM inference into efficient streaming reconstruction by exploiting this temporal redundancy. Our key observation is that, under fixed-view capture, representations of static scene content can be effectively reused across timestamps, while only dynamic regions require continuous updates. Accordingly, \ourname decouples temporal computation into static reuse and dynamic token updates: static representations are carried forward from the reference frame, whereas dynamic tokens are selectively recomputed from current observations. This principle is independent of a specific interaction mechanism and can be readily applied to both attention- and TTT-based reconstruction models without additional temporal modules or retraining. Extensive experiments on diverse multi-view dynamic benchmarks demonstrate substantial inference acceleration with minimal degradation in reconstruction quality, establishing \ourname as a simple and effective framework for streaming feed-forward volumetric video reconstruction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.