RDA: Recovering Dropped Contributions from Retained Attention for Video Generation
Abstract
The quadratic cost of self-attention is a major bottleneck in video diffusion transformers. Training-free sparse attention with online top- selection has emerged as a flexible solution, but existing methods struggle to reconcile mask adaptivity with execution efficiency. Block-wise methods enable efficient regular computation, yet their coarse partitions limit the flexibility of the resulting masks; centroid-wise methods construct more content-adaptive masks, but incur costly preprocessing and irregular computation. We propose RDA, a training-free sparse attention method that retains the efficiency of regular block-wise execution while recovering information lost by coarse block masks. RDA co-designs recovery-aware top- block selection and compensation to predict dropped contributions from already-computed retained attention outputs. We further introduce a sensitivity-driven two-dimensional scheduler that jointly assigns top- configurations across denoising steps and network layers, minimizing a calibrated attention-error objective under a target sparsity budget. On 720p HunyuanVideo and Wan2.1, RDA achieves 2.010 and 1.663 speedups at 78.3% and 77.1% sparsity, respectively, while providing the best fidelity among sparse baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.