Temporal Residual for Consistent Autoregressive Long Video Diffusion
Abstract
Autoregressive video diffusion enables efficient long-video generation, but maintaining visual consistency while allowing scenes to evolve remains challenging. Existing approaches improve robustness to generated histories or retain selected visual anchors, yet bounded context windows lose access to earlier appearance cues, while fixed visual anchors cannot summarize an evolving history and restrict video motion. To resolve the three-way tension among long-range consistency, meaningful dynamics, and the efficiency of bounded-context autoregressive video generation, an efficient global memory that evolves with video progress is needed alongside local context and fixed visual anchors. We introduce Temporal Residual (TR), a lightweight evolving plug-in global state that extends pretrained generators through brief post-training. TR combines three key designs: Evolving Global State, Dual-Path Memory Injection, and Chunk-Synchronous Evolution Schedule. Together, they make evolving visual history available throughout generation while allowing the scene to continue changing. Applied to Self Forcing and LongLive with brief post-training, TR improves long-range consistency and visual quality while preserving meaningful dynamics. Human preferences and ablations further support the value of this global state.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.