acceptodds
Under review as a conference paper at ICLR 2027

Trace Forcing: Anchor Tokens in the KV Cache for Training-Free Long Video Generation

Abstract

Autoregressive video diffusion models generate videos of arbitrary length by conditioning each frame on the preceding frames stored in a KV cache, yet their quality degrades as generation proceeds beyond the training horizon (the video length seen during training). Existing training-free methods attribute this degradation to RoPE temporal positions beyond the trained range and re-index the temporal positions of the cached tokens at inference, but the degradation persists. We reveal that the cached tokens themselves deviate from the tokens the model was trained on, and that the deviated tokens entering the cache induce artifacts in the regions that attend to these tokens. To address this, we introduce Trace Forcing, a training-free method that replaces the raw tokens in the KV cache with anchor tokens, which aggregate the corresponding tokens across frames by an exponential moving average along traces obtained from the similarity between adjacent frames. Each anchor token thus incorporates earlier representations that are less affected by the deviation while remaining consistent with the corresponding content, so that the model conditions on representations closer to the tokens observed during training. We further propose adaptive updating, which assigns each anchor an update rate according to how its content evolves, suppressing the accumulation of deviation on static content while allowing dynamic content to follow recent information. Experiments show that Trace Forcing exhibits less quality degradation than existing training-free methods when generating far beyond the training horizon.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.