acceptodds
Under review as a conference paper at ICLR 2027

Linear Indexer: Understanding and Accelerating DeepSeek Sparse Attention

Abstract

DeepSeek's Lightning Indexer scores every query against all visible keys, incurring quadratic prefill cost. In our DeepSeek-V4-Flash profile at 256K context, it accounts for % of prefill GPU busy time. Surprisingly, we find that removing its head-wise ReLUs largely preserves retrieval. This observation yields **Linear Indexer**, a training-free method that folds the 64 weighted query heads exactly into one before the global scan, reducing 64 query–key dot products per key to one. The resulting linear score directly selects keys for attention; for higher retrieval recall, the original scorer can rerank an enlarged linear shortlist. Our analysis links the high recall to learned coupling between stable query–key signs and head weights, a pattern observed across domains, context lengths, checkpoints, and Indexer architectures. On V4-Flash, Linear Indexer maintains performance comparable to the original model on LongBench-v2 and RULER at context lengths up to 256K. At 256K on NVIDIA H20 GPUs, direct mode achieves a indexer speedup and a speedup in end-to-end prefill time to first token (TTFT).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.