acceptodds
Under review as a conference paper at ICLR 2027

Recency Forcing: Scaffolding Temporal Dependency for Long Autoregressive Video Generation

Abstract

Autoregressive (AR) video generation is causal by construction, yet how a generated frame depends on individual past context frames remains poorly understood, especially across the denoising steps of few-step distilled models. We introduce the positional response , a perturbation-based sensitivity measure quantifying how strongly predictions depend on a context frame at temporal distance and denoising step . Measured across distilled AR video diffusion models, reveals a consistent two-axis structure: influence decays rapidly with distance, while later denoising steps draw on a broader temporal window than earlier ones. Motivated by this, we propose Recency Forcing, applying a timestep-dependent Temporal Response Bias (TRB) to attention logits that scaffolds attention to match the measured response. This also mitigates KV eviction mismatch, where distant frames must be evicted from the KV cache during long-horizon inference despite being available in training: since TRB already drives their influence toward zero, eviction has little impact, letting short-clip-trained models generate substantially longer videos. We further introduce Biased Attention Reparameterization (BAR), an exact reformulation compatible with FlashAttention at negligible overhead. Recency Forcing supports both training-free and training-based settings, and experiments on VBench and VBench-Long show consistent long-horizon quality improvements with negligible inference overhead.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.