acceptodds
Under review as a conference paper at ICLR 2027

When the Future Remembers: Past Reconstruction in Autoregressive Long Video Generation

Abstract

As autoregressive video diffusion models become the new paradigm for long-form video generation, a fundamental issue remains: these models tend to forget earlier content once it falls outside the bounded KV cache, resulting in identity drift and scene resets that compromise long-term temporal consistency. We propose Eviction Forcing, a simple yet effective training strategy that teaches the model to recover the past by conditioning on its own future. After evicting a prefix of frames from a standard base teacher rollout, a lightweight LoRA adapter is trained to reconstruct the evicted frames at their original RoPE indices, using only the remaining frames that are later on the temporal axis as context. This approach teaches the model to remember the past beyond the immediate attention window, thereby significantly improving video consistency across long durations. The model is further optimized via periodic supervision from the base teacher model under full-context rollouts, ensuring its stable performance during standard rollouts. Extensive experiments demonstrate that Eviction Forcing achieves competitive overall VBench scores and consistently improves long-range temporal consistency over state-of-the-art autoregressive baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.