Local Forcing: Local KV-Cache Construction for Long Video Generation
Abstract
One of the central challenges in causal video generation is error accumulation. We find that attending to history during key–value (KV) cache construction may exacerbate error accumulation, even though the same context is essential for denoising. This observation motivates Local Forcing (LF), which follows a simple principle: generate with context; encode locally. This principle is implemented through an attention-mask modification applied consistently across teacher forcing, consistency distillation, self-rollout distillation, and inference. Surprisingly, LF improves robustness to drift, enabling stable minute-level rollouts without specialized attention-sink strategies, additional memory modules, or long-horizon distribution matching distillation. The simplicity of LF facilitates integration with both training-aware and training-free approaches to long-video generation, yielding consistent improvements in overall video quality across the evaluated methods. These findings position local cache construction as a foundational design principle for scaling causal video generation far beyond its training horizon.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.