Ms. Forcing: Efficient Streaming Video Generation with Multi-Scale Patchification and Attention
Abstract
Streaming video diffusion models have made substantial progress toward interactive and dynamic world simulation, but the nested autoregressive and denoising loops of conventional next-frame generation hinder real-time deployment. Recent rolling-window methods pipeline denoising across multiple consecutive frames at different noise levels, improving throughput and long-horizon stability. However, they tokenize every state at the same fine spatial granularity, leaving substantial noise-dependent redundancy in the joint denoising window. We propose Ms. Forcing, which adapts spatial granularity to each state's noise level for efficient streaming video generation. Its Multi-Scale Patchification (MSP) assigns coarser patches to noisier states, reducing the active-window token count by 45%, while Multi-Scale Self-Attention (MSSA) matches the density of visible non-sink keys and values to each query scale to further reduce attention cost. Fixing both schedules by window position preserves a static, hardware-friendly computation graph in Ms. Forcing. We further reduce the training–inference mismatch in rolling-window DMD through Homogeneous-Noise-Level DMD (H-DMD), which assembles each fake video from clean predictions sharing the same source noise level. The multi-scale design helps offset the additional training cost of backpropagating through overlapping windows. We include both quantitative and qualitative experiments to show that Ms. Forcing reaches 22.84 FPS on a single H200 GPU, 39.7% faster than Rolling Forcing, while significantly improving VBench scores in both short video and long video generation settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.