acceptodds
Under review as a conference paper at ICLR 2027

Grid Compressed Attention: Accelerating Video Generation Attention with Training-Free Grid KV Compression

Abstract

Attention computation is a major bottleneck in video Diffusion Transformers because its cost grows quadratically with the sequence length. Existing training-free sparse attention methods reduce computation via sparse token selection, but the more evenly distributed attention mass in video generation limits how aggressively tokens can be discarded without losing information. We introduce Grid Compressed Attention (GCA), a training-free method that compresses attention via KV compression and query grouping. Specifically, we formulate attention compression as an entropy-regularized optimization problem and derive an efficient construction based on compact KV compressions and disjoint query groups. A local Fisher characterization of the group-logit approximation error guides the selection of representative queries. Experiments on Wan2.1-14B show favorable fidelity–latency trade-offs for GCA among the evaluated configurations. These results support query-guided KV compression as a promising training-free approach to accelerating video generation attention.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.