acceptodds
Under review as a conference paper at ICLR 2027

Clean Forcing: Drift-Resistant Autoregressive Video Diffusion with a Frozen Base

Abstract

Autoregressive (AR) video diffusion models generate streaming video of arbitrary length by conditioning each new chunk of frames on those already generated. However, they suffer from exposure bias, where a model trained on clean context must continue from its own imperfect outputs, and errors compound until quality collapses within seconds (drift). Existing methods either require cluster-scale base retraining, apply training-free inference corrections with inconsistent results, or train frozen-base correctors for sampling and cache errors. We instead find that drift is largely deterministic and therefore learnable, as holding the noisy state fixed while varying only the history reveals a counterfactual velocity gap that is 95% systematic across noise seeds. Based on this observation, we introduce *Clean Forcing*, which trains a 5.9M-parameter LoRA corrector on the *frozen* base model to regress this gap, with a closed-loop stage that adds DAgger-style aggregation and a drift-contraction objective to expose the corrector to its own rollouts. Because the clean histories can be the base model's own single-shot generations, training needs *no real videos*, and the merged corrector adds no inference cost. In 50s text-to-video generation, Clean Forcing reduces -drift from to without real video and to with 40 real clips, improving on the best published result of with over less data. Extensive experiments demonstrate that Clean Forcing outperforms prior methods in aesthetic quality and data efficiency with reduced drift for long-horizon video generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.