acceptodds
Under review as a conference paper at ICLR 2027

Accelerating Causal Video Generation with Mixed Sparse Attention

Abstract

Causal video diffusion has emerged as a streaming alternative to bidirectional models that denoise entire clips before playback. These autoregressive models use a key/value cache, but attention costs grow with visible history. However, causal masking changes attention support, so directly adapting bidirectional sparse attention can degrade quality. Prior work also leaves attention patterns across chunks, denoising passes, and horizons underexplored. We find stable spatial-local heads, weakening recent-window support, and changing reads despite immutable cached keys. Therefore, Mixed Sparse Attention (MSA) combines static/content routing, horizon/cost gates, reusable statistics, and budgeting across passes. For 20-second Causal Forcing generation at 720p, MSA achieves up to 2.634 denoising and 1.550 end-to-end speedup over dense attention. Moreover, MSA achieves comparable VBench and VBench-Long scores to the matched sparse-attention baseline across six evaluated dimensions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.