Accelerating Causal Video Generation with Mixed Sparse Attention
Abstract
Causal video diffusion has emerged as a streaming alternative to bidirectional models that denoise entire clips before playback. These autoregressive models use a key/value cache, but attention costs grow with visible history. However, causal masking changes attention support, so directly adapting bidirectional sparse attention can degrade quality. Prior work also leaves attention patterns across chunks, denoising passes, and horizons underexplored. We find stable spatial-local heads, weakening recent-window support, and changing reads despite immutable cached keys. Therefore, Mixed Sparse Attention (MSA) combines static/content routing, horizon/cost gates, reusable statistics, and budgeting across passes. For 20-second Causal Forcing generation at 720p, MSA achieves up to 2.634 denoising and 1.550 end-to-end speedup over dense attention. Moreover, MSA achieves comparable VBench and VBench-Long scores to the matched sparse-attention baseline across six evaluated dimensions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.