HyMoR-Attn: Granularity-Matched Repair of Omitted Attention Support for Video Diffusion
Abstract
Block‑sparse attention has become the backbone of efficient video diffusion. A growing body of work seeks better fidelity and efficiency trade‑offs by improving the accuracy of coarse block selection. Intuitively, a more accurate selector should yield a better frontier. However, we find that coarse‑only selection can itself limit the attainable frontier. Video‑diffusion attention is inherently two‑scale. Coarse blocks efficiently capture a locally dense body while a scattered tail of high‑value interactions remains inside omitted blocks. We call this failure the *Omitted‑Micro Problem*. We further find that omission risk is hierarchically predictable. Risk is stable at the layer and head level and stable enough for offline per‑model calibration while the coordinates of valuable Microtiles remain content‑dependent and require online selection. This observation motivates a rethink of selection‑centric sparse attention. Improving coarse rankings alone cannot recover fine interactions that the execution granularity makes invisible. We introduce **H**ybrid **M**icro **O**mission **R**epair (HyMoR‑Attn), a training‑free support‑repair method built on a coarse Macro backbone with a fine‑grained Micro residual. It combines offline risk calibration to locate vulnerable layers and heads, online peak‑preserving selection to recover the omitted tail, and granularity‑matched Macro and Micro execution with exact LSE merging. Across four settings spanning two models, two tasks and 480p and 720p resolutions, HyMoR‑Attn is Pareto‑optimal among the compared methods. It achieves up to 2.06 and 2.12 speedup over Full Attention while reaching up to 33.01 and 28.46 PSNR on HunyuanVideo and Wan-2.1 respectively. The quality ceiling of coarse block sparsity is set not by the chosen budget but by the high‑value support it discards. Matching repair granularity to local interaction density raises fidelity and efficiency together.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.