acceptodds
Under review as a conference paper at ICLR 2027

Residual Forcing: Structure-Aware Sparse Attention for Autoregressive Video Diffusion

Abstract

Autoregressive video diffusion enables streaming long-video generation, but self-attention over cached visual history remains a major inference bottleneck. Training-free block-sparse methods can flexibly select key blocks for each query block, yet routing based on inexpensive mean-pooled scores can miss concentrated interactions amplified by softmax. These misses degrade generation quality, while compensating with more key blocks erodes the efficiency gains; independently changing the selected context across chunks can also introduce periodic visual discontinuities. In this paper, we show that many of the missed high-response interactions follow identifiable attention structures that can be protected at low cost. We therefore introduce , a training-free framework that separates structure preservation from residual attention selection. It protects spatially aligned key blocks, their axial halo tokens, and shared salient keys, then routes only the residual key space using reusable head-level priors and block-adaptive budgets. To improve temporal continuity, it shares candidate-frame choices across chunk boundaries without increasing the frame budget. Fused routing and dynamic sparse-attention kernels efficiently execute the resulting heterogeneous routes. On LongLive, Residual Forcing achieves a self-attention speedup in fixed-cache profiling and a end-to-end speedup for 30-second generation, while maintaining competitive video quality and mitigating chunk-boundary discontinuities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.