acceptodds
Under review as a conference paper at ICLR 2027

Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models

Abstract

Streaming autoregressive diffusion models enable real-time, long-horizon video generation, but their training objectives optimize local frame prediction rather than the geometry and dynamics of a coherent world: long rollouts accumulate geometric drift and degrade into static or unnatural motion. Recent bidirectional approaches address this problem using rewards signals built upon 3D Gaussian-Splatting reconstruction. However, a single rigid 3d reconstruction cannot model a dynamic scene, so this critic *penalizes* genuine object motion as reconstruction error and is maximized by freezing the video. This shortcut is especially detrimental in the AR setting, where each chunk can propagate an already-static configuration. In this work, we propose Stream4D, which replaces the static critic with a feed-forward 4D reconstruction reward that explicitly models scene dynamics, allowing coherent motion to receive high consistency rewards. To further guide motion magnitude and quality, we add a motion prior that rewards natural scene-flow magnitude while penalizing jitter and non-rigid artifacts. Our final recipe combines these two terms with a lightweight perceptual anchor. On the 5-second Self-Forcing and Causal-Forcing backbones and the 10-second LongLive backbone, Stream4D improves 4D reconstruction PSNR by 3.46, 5.53, and 6.76 dB, respectively, while better preserving motion dynamics and earning higher human preference. Videos can be found in the supplementary material.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.