acceptodds
Under review as a conference paper at ICLR 2027

Temporal Residual for Consistent Autoregressive Long Video Diffusion

Abstract

Autoregressive video diffusion enables efficient long-video generation, but maintaining visual consistency while allowing scenes to evolve remains challenging. Existing approaches improve robustness to generated histories or retain selected visual anchors, yet bounded context windows lose access to earlier appearance cues, while fixed visual anchors cannot summarize an evolving history and restrict video motion. To resolve the three-way tension among long-range consistency, meaningful dynamics, and the efficiency of bounded-context autoregressive video generation, an efficient global memory that evolves with video progress is needed alongside local context and fixed visual anchors. We introduce Temporal Residual (TR), a lightweight evolving plug-in global state that extends pretrained generators through brief post-training. TR combines three key designs: Evolving Global State, Dual-Path Memory Injection, and Chunk-Synchronous Evolution Schedule. Together, they make evolving visual history available throughout generation while allowing the scene to continue changing. Applied to Self Forcing and LongLive with brief post-training, TR improves long-range consistency and visual quality while preserving meaningful dynamics. Human preferences and ablations further support the value of this global state.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.