acceptodds
Under review as a conference paper at ICLR 2027

Token-Gated Sparse Attention

Abstract

Attention-based architectures struggle in long-context decode settings due to the need to load and process the full set of attention keys and values for the sequence (the KV cache) for each newly generated token. Sparse attention approaches alleviate this bottleneck by computing attention over a subset of the KV cache. Prior sparse attention approaches generally impose restrictive structure on the sparsity pattern, such as targeting a fixed amount of sparsity within a given attention head or at any given context length. We propose token-gated sparse attention (TGSA), a new variant of sparse attention that allows each token within a given attention head to choose which elements of the KV cache it reads, resulting in a dynamic and non-uniform allocation of sparsity across heads, layers, and contexts. TGSA introduces an auxiliary set of small gating heads, which control whether or not a given element of the regular attention KV cache must be loaded. Computing the gating heads before regular attention allows TGSA to skip chunks of the (much larger) standard attention KV cache load. To evaluate TGSA, we pretrain models at 600M, 1.7B and 4B scales and find that its textual quality is comparable to regular attention and post-trained sparse attention baselines, and that TGSA leads to up to 3.5x faster decoding than regular attention at 64K context length.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.