AR: Training-Free Sparse Attention for Autoregressive Video Diffusion
Abstract
Autoregressive video diffusion has recently emerged as a promising paradigm for real-time and streaming video generation, owing to its causal decoding formulation and natural support for key–value (KV) caching. However, as generation proceeds, the KV cache grows monotonically over rollout, making attention over long KV histories the dominant inference bottleneck. In this work, we propose \textbf{A^2R} (nchored and daptive eusable Attention), a training-free sparse attention framework for long-horizon autoregressive video diffusion. Our design is motivated by three observations: early chunks act as persistent visual anchors for later generation, attention sparsity varies substantially across heads and routing patterns remain stable across intermediate denoising steps. Based on these observations, AR combines an anchored dense-to-sparse transition, per-head adaptive block routing and reusable sparse patterns that are refreshed only at the final denoising step. We further implement an exact Triton-based sparse attention kernel that computes attention only over retained KV blocks. Extensive experiments on autoregressive video diffusion backbones show that AR achieves around 77% average attention sparsity, reduces attention runtime by \textbf{2.1\times}, and delivers a \textbf{1.33\times} end-to-end speedup while preserving strong visual quality and temporal consistency, yielding a favorable quality–efficiency trade-off for video generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.