ResiSparse: Residual-Informed Exact-or-Approximate Sparse Attention for Efficient Video Generation
Abstract
Video diffusion transformers repeatedly apply spatiotemporal self-attention, making denoising expensive as video sequences grow. Training-free sparse attention reduces computation, but discarding unselected interactions can make model outputs deviate from their dense-attention counterparts. A recent paradigm computes selected interactions exactly and approximates the rest. Accurate approximation remains essential: replacing individual keys and values with cluster centroids can still distort the output. We introduce ResiSparse, a training-free sparse-attention method guided by a theoretical analysis of centroid-based approximation. This analysis motivates three components: (i) Query-Induced Key Clustering uses query information to reweight key channels before clustering, reducing centroid-induced logit distortion; (ii) Unified Mass–Dispersion Routing combines centroid logits with within-cluster key and value variation to prioritize exact computation using a cost-normalized score; and (iii) Query-Prioritized Taylor Compensation uses a second-order expansion to account for logit variation within approximated blocks, correcting their centroid-based contributions to both the attention numerator and softmax denominator. Across Wan2.1-1.3B, Wan2.1-14B, and HunyuanVideo-13B, ResiSparse outperforms the evaluated training-free sparse-attention methods and achieves a stronger quality–efficiency trade-off, with PSNRs of , , and at , , and denoising speedups, respectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.