acceptodds
Under review as a conference paper at ICLR 2027

BRAID: Band-Gate Routed Attention with Interleaved Detail for Efficient Video Diffusion Transformers

Abstract

Exact self-attention is the computational bottleneck of video diffusion transformers: its cost grows quadratically with the full spatiotemporal token sequence, while every other block in the network scales linearly. We present Band-Gate Routed Attention with Interleaved Detail (BRAID), a sparse-linear attention module that replaces dense self-attention with two complementary paths: a query-adaptive block-sparse branch that evaluates exact softmax on selected key blocks under a fixed budget; and a subband-gated linear complement that carries global context across Haar wavelet subbands through learned per-head gains. A structural sink reservation stabilizes the sparse normalizer, and a denoise-progress-aware schedule reallocates one average block budget across the sampling trajectory. BRAID is defined at the attention-module level and is not tied to a particular backbone. We establish a unified cost decomposition across Dense as an upper-bound reference, SLA as the direct sparse predecessor, and BRAID, keeping this arithmetic analysis separate from matched measurement. On the full evaluation set, BRAID improves prompt following over its direct sparse predecessor SLA while remaining within a narrow margin of the exact-attention reference; on the reference hardware, one BRAID forward step takes 46.5% less wall-clock time than the exact-attention step of the same network.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.