acceptodds
Under review as a conference paper at ICLR 2027

Native Token-Level Sparse Attention for Long-Context Video Pretraining

Abstract

Modern video generation models built on bidirectional diffusion transformers (DiT) run quadratic attention over space–time tokens. This O(N²) cost bottlenecks the scaling of video DiT to longer and higher-resolution generation. Sparse attention mitigates the cost by limiting each query to a subset of key–value pairs. Existing sparse attention mechanisms for video DiT mainly apply block-wise selection, which keeps the implementation hardware-friendly and respects the spatio-temporal locality of nearby video tokens. However, block-wise selection is coarse: whenever the relevant context is finer than a block, the attention budget is spent on irrelevant tokens while relevant ones are missed. The resulting gap to full attention is acceptable for accelerating inference, but not for a native sparse attention. It is the attention of the model throughout training, so it pays this gap at every update and has to match full attention. In this paper, we propose VTSA (Video Token-level Sparse Attention), a native token-level sparse attention that is lossless in long-context video pretraining. VTSA selects individual key–value tokens with a small indexer distilled from the model's own attention scores. One selection is shared across the query heads that read one key–value head and across a small tile of neighbouring queries. A custom kernel gathers the selected tokens in the forward pass and scatters their gradients in the backward pass. We compare VTSA with two trained sparse-attention baselines, VSA-style block routing and sliding-window attention, in continued pretraining from dense checkpoints. We use two video DiT backbones with different attention architectures, a 5B multi-head video model (Wan2.2-5B) and a 15B grouped-query audio–video model (daVinci-MagiHuman). VTSA matches full attention while each sparse layer attends to 6–15% of the key–value pairs, at 16k and 32k tokens on the 5B backbone and from 16k to 128k tokens on the 15B backbone. Block routing matches full attention only at two to four times as many tokens. VTSA opens the door to long-context video pretraining with native token-level sparse attention.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.