acceptodds
Under review as a conference paper at ICLR 2027

FlashIndexer: Dynamic Head-Routed Indexer for Token-Level Sparse Attention

Abstract

Token-level sparse attention, exemplified by DeepSeek Sparse Attention (DSA), significantly reduces the downstream attention computation, but shifts the computational burden to the lightweight indexer. For every query, the indexer must compute scores for all context tokens using its entire set of heads before selecting the most relevant top- tokens, resulting in a per-layer cost that grows with both the context length and the number of indexer heads. Recent work has reduced indexer computation along the token and layer axes, but the head axis has received comparatively little attention. We revisit the head axis and make two key observations. First, the contributions of indexer heads are highly imbalanced and vary substantially across samples. Second, the important heads for each query can be reliably identified using only the query itself, without accessing any context tokens. Motivated by these observations, we propose FlashIndexer, a plug-and-play head-routing module for the DSA indexer. For each query, it dynamically selects a small subset of important heads from all indexer heads, activating only the selected heads for scoring context tokens, and thereby substantially reducing indexer computation. This query-adaptive routing preserves the original token-level top- selection while avoiding unnecessary computation from less informative heads. With only a small subset of indexer heads activated, FlashIndexer matches the accuracy of the original full-head DSA indexer on DeepSeek-V4-Flash-0731 and GLM-5.1 across RULER and LongBench v2, while achieving up to end-to-end time-to-first-token speedup at K context.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.