acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Sparse Attention in Few-Step Diffusion Transformer

Abstract

Few-step diffusion compresses a multi-step teacher trajectory into a handful of student evaluations. Sparse attention further reduces cost by retaining only highly ranked interactions, but the student’s current ranking need not represent the corresponding teacher-window support. On aligned probes of a four-step Wan2.2 student, the top-b pool captures 51.91% of teacher attention mass and leaves a value-weighted residual. The useful core favors persistent interactions; its complement contains both missed salient blocks and a diffuse tail. This diagnosis motivates SPARE Attention: current-score candidates plus a small complement quota under a fixed sparse budget, with coarse routing and no teacher access at inference. A matched per-step-budget intervention reduces final Teacher40 RMSE from 0.959537 to 0.928340 (3.25%), improving all six reported records; larger quotas are not uniformly better. Separately, the original system configuration selects 14.84% of key blocks on the full 6,220-video VBench suite and denoises in 29.565 s, the shortest measured time among the timed sparse implementations, with Total comparable to the SLA family and PISA

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.