acceptodds
Under review as a conference paper at ICLR 2027

GroupIndex: Accelerating Sparse Attention via Learned Cross-Query Indexing

Abstract

DeepSeek Sparse Attention (DSA) makes attention cheap by keeping only a few keys per query, but its lightweight indexer still scans the whole prefix for every query, and this scan comes to dominate long-context prefill. Existing accelerators reuse selections across layers, activate fewer indexer heads or filter key blocks. Sharing across layers and heads is capped by the fixed numbers of indexer layers and heads, and accuracy grows harder to keep as it nears that cap. Filtering key blocks instead trades the resolution of their summaries against recall. Sharing one scan across neighboring queries has no such fixed cap, since how many can share depends on their similarity and on context length. Existing query-group indexing, however, averages the queries of each group and loses accuracy as the group widens. We show that this width limit is representational rather than a matter of capacity. A shared scan must keep every key that any member selects, which calls for the maximum of the members' scores, whereas averaging their queries dilutes each member's own needs and loses keys that are important to individual members. We introduce GroupIndex, which jointly learns head-wise query aggregation and scoring projections in a Surrogate Indexer, and applies hardware-aware layer specialization to select each layer's query-group width and candidate budget. It indexes each query group with one scan whose surrogate query is learned rather than averaged. Training targets the group's selection union, uses the native indexer's own selections as labels, and runs offline on activations cached once, without loading the backbone. At inference the Surrogate Indexer runs as a sidecar ahead of the native indexer and proposes candidates for each group, among which the native indexer still makes each query's final selection. On DeepSeek-V4-Flash, GroupIndex matches the dense indexer's accuracy on RULER, LongBench-v2 and LongMemEval from 64K to 512K tokens, and at 1M tokens it speeds up the indexer about 11 and end-to-end prefill up to 1.9.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.