Wave Forcing: Towards Speed-of-Light Streaming Video Generation
Abstract
Real-time long-video generation requires fast, continuous output with consistent appearance and natural motion. Chunk-wise autoregressive methods serialize denoising and history-cache refresh, requiring two forwards even with one-step denoising. Their long rollouts also often suffer from quality degradation and camera drift. We propose **Wave Forcing (WF)**, a co-design of generation, training, and execution for continuous output with full multi-step denoising. WF pipelines chunks along a causal wavefront at staggered denoising stages, using chunk-wise causal attention to access earlier, cleaner chunks before they finish. Distributing denoising and cache-refresh stages across full-model replicas yields **1F1C**: one chunk per concurrent forward round after pipeline fill. Three-stage training develops joint mixed-noise denoising, strengthens motion and framing, and adapts the model to causal inference with self-generated history. **Wave Runtime (WaveRT)** uses **Wave Parallelism (WP)** for early K/V delivery and communication–computation overlap, complemented by kernel optimization and load-balanced VAE decoding. On Wan2.1-T2V-1.3B, 5-step WF maintains overall quality comparable to Rolling Forcing with higher Dynamic Degree scores on short and long videos. With SageAttention, our 4-step model reaches 117.7 E2E FPS on eight H200 GPUs, delivering a speedup over the 4-step single-GPU baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.