QUASAR: Query-Adaptive Sparse Attention with Recovery for Video Generation
Abstract
Long spatiotemporal attention in video Diffusion Transformers becomes an inference bottleneck as sequences grow. Existing sparse attention methods can be tightly coupled to sparse-operator implementations, while accuracy compensation may be insufficient or overly complex, limiting the speed-accuracy trade-off of sparse attention. We propose QUASAR, which uses a block-relevance score to select KV blocks for exact computation and restores the contributions of the remaining blocks. Selection has two variants: Full-Domain scores all legal KV blocks, while Candidate-Constrained applies the same score within a candidate set provided by an external sparse-attention method. Recovery estimates the contributions of unselected blocks from Value-block summaries and combines them with exact attention under the same normalization, with Recovery states at multiple query granularities. We prove output-error bounds for both Candidate-Constrained and Recovery. In Wan2.1, QUASAR with per-query Recovery at sparsity achieves VBench AVG5 of (Dense: ) with end-to-end speedup. At 720p, QUASAR reaches up to end-to-end speedup with more than sparsity and remains about VBench AVG5 drop relative to dense attention. Results across sparse-attention backends, video models, KV-block geometries, and hardware demonstrate QUASAR's portability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.