Stream3D: Sequential Multi-View 3D Generation via Evidential Memory
Abstract
View-conditioned 3D generators such as SAM3D, TRELLIS and Hunyuan3D produce high-quality object 3D representations from a single view, but real-world visual observation often arrives as long monocular streams. Each frame provides a partial observation; views accumulated across the stream provide complementary evidence for reconstruction. To address this problem, we propose Stream3D, the first training-free streaming mechanism that turns a frozen view-conditioned 3D generator into a streaming generator with constant cross-chunk evidential memory. Stream3D achieves this by maintaining a compact evidential memory, which selectively caches the most informative historical frames based on a proposed evidence score mechanism. As the stream progresses, the memory dynamically updates to retain a fixed number of informative frames, preventing the evidence cache from growing linearly with sequence length. This sustains generation quality over long sequences and keeps the underlying generator completely unchanged without retraining, architectural modifications, or auxiliary losses. Evaluated on both realistic and synthetic streaming benchmarks, Stream3D outperforms latent-transport baselines, including KV-cache reuse and flow-based feature editing, across both photometric and geometric metrics. More details can be found at: [Link](https://anonymous-submission-20.github.io/streaming3D2.github.io/)
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.