MovieRoll: Rolling Out Long-Form Videos with Hierarchical KV Caching and Cumulative Long Tuning
Abstract
Streaming video generation seeks to extend high-quality short-video synthesis to continuous, interactive generation. However, the mismatch between training on short videos with ground-truth context and long-horizon inference conditioned on generated content makes quality difficult to sustain. Meanwhile, full-history KV caching incurs memory costs that grow with generation length. To address these challenges, we present MovieRoll, a base-to-distilled framework that first establishes a multi-step streaming base through supervised training and then distills it into a few-step generator through long self-rollouts. At the base stage, Window-wise Teacher Forcing jointly denoises multiple chunks with progressive noise levels within each window, preserving within-window temporal interactions and improving resistance to long-horizon degradation compared with chunk-wise training. To retain historical context efficiently, hierarchical KV-cache compression (HieraKV) exploits layer-wise attention sparsity in streaming DiTs, retaining the global available history in shallow layers but only the sink chunk and the most recent history chunk in deep layers, which reduces KV-cache memory by 78% and improves degradation resistance. Building on this base, cumulative long tuning with distribution matching distillation (CuLT-DMD) accumulates and averages gradients from different local clips throughout an entire long self-rollout. The resulting gradient spans the full sample and provides a rollout-level estimate, improving stability during long-horizon generation. Finally, at inference time, threshold-triggered history selection and KV recomputation bound the retained context, while classifier-free guidance annealing (AnnealCFG) progressively reduces guidance strength toward the clean endpoint during denoising to further stabilize MovieRoll-Base. In 30-second HeliosBench evaluations, MovieRoll-Base achieves competitive generation quality. MovieRoll-Distilled achieves state-of-the-art (SOTA) overall scores in both single- and multi-event settings, combining the highest naturalness with sufficient dynamics. Model weights and code will be publicly released after the review process.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.