acceptodds
Under review as a conference paper at ICLR 2027

In-Distribution Forcing for Long Video Generation at Test-Time

Abstract

While modern autoregressive (AR) video diffusion models have demonstrated remarkable success in short-horizon video generation, generating long videos remains challenging due to a phenomenon known as drifting, where colors and textures shift, and motion dynamics decay. Existing approaches primarily rely on KV conditioning, which selects or modifies cached key-value (KV) entries to mitigate drifting. However, we observe that KV conditioning alone is insufficient because it assumes cached KV entries remain in-distribution. This assumption fails beyond the training horizon because no mechanism constrains the construction of cached KV entries during rollout, giving rise to the KV-provenance problem where cached entries themselves become out-of-distribution. To address this limitation, we propose In-Distribution Forcing (ID-Forcing), a test-time framework that aligns both KV caching and KV conditioning with training configurations. Its key mechanism, self-caching, prevents out-of-distribution KV entries at their source: each chunk is cached without attending to any prior KV entry, keeping the rolling window exactly in-distribution without retraining. Consequently, ID-Forcing seamlessly extends short-horizon models to minute-scale long video generation. Extensive evaluations show that our method remains competitive on standard video generation benchmarks while substantially outperforming prior work in mitigating drifting, as validated by both our drift metrics and a user study.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.