LISA: Head-Fused Linear Indexing for Efficient Sparse Attention
Abstract
Sparse attention reduces the cost of long-context inference by restricting attention to a small subset of tokens for each query. However, identifying this subset can still be expensive: the indexer in DeepSeek Sparse Attention (DSA) scores every visible token with multiple query heads, leaving substantial indexing overhead during long-context prefilling. We introduce LISA, a training-free indexer based on an exact decomposition of the weighted-ReLU score into a linear term and a nonlinear correction. Shared keys allow the linear contributions of all heads to be fused exactly into one dot product per token. LISA directly selects tokens using this linear term, omitting the nonlinear correction and removing the head-count factor from full-context scoring while preserving token-level granularity. LISA adds candidate refinement with the original multi-head scorer, while LISA shares candidate pools across layers with layer-specific refinement. Across 4K–128K contexts, LISA maintains RULER scores close to the respective native baselines on DeepSeek-V3.2 and GLM-5.2. On DeepSeek-V3.2 with a 1M-token input, LISA achieves a speedup over DSA in single-request offline time to first token, while LISA with four-layer groups achieves . These results support head-fused linear scoring as an efficient approximation to multi-head indexing for long-context prefilling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.