acceptodds
Under review as a conference paper at ICLR 2027

Step Forcing: Training-free Step-Matched KV Caching for Long Video Generation

Abstract

Autoregressive video diffusion models adapt bidirectional video diffusion models into causal architectures, enabling long video generation through autoregressive rollout. However, these models often suffer from model drift over time, leading to progressive degradation in visual quality. Existing studies generally attribute such model drift to error accumulation during autoregressive rollout. Although existing methods alleviate error accumulation, they overlook a fundamental mismatch in the chunk-wise diffusion process, where the noisy state of the current chunk attends to the same KV caches extracted from previously generated clean chunks. To address this mismatch, we present Step Forcing, a training-free approach designed to mitigate model drift during autoregressive long video generation. We find that reusing fixed clean KV caches throughout the denoising trajectory introduces a step-wise representational mismatch and that noisy and clean KV caches serve complementary roles in preserving visual quality and temporal consistency. Building on these insights, Step Forcing reuses noisy KV caches from corresponding denoising steps while retaining a subset of clean anchor tokens. Our proposed Step Forcing effectively mitigates model drift by using step-matched KV tokens, enabling baseline models to generate long videos with high visual quality and temporal consistency. Results on benchmark quantitatively demonstrate that Step Forcing improves the performance of baseline models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.