acceptodds
Under review as a conference paper at ICLR 2027

DoSA: Don't Score All, Certified Sparse Attention Indexing

Abstract

Sparse attention visits few keys, yet finding them can still require scoring every key with every indexer head, leaving a dense computation bottleneck inside the attention module. We introduce DoSA (Don't Score All), a lossless optimization that makes token selection conditional while preserving exact Top- results. Its central observation is that rejecting a candidate requires a sufficiently tight score bound, rather than its full score. A seed set supplies a lower bound on the selection threshold. Nonnegative head projections and joint residual bounds upper-bound positive contributions, while activation-informed aggregation bounds the magnitude subtracted by negative heads. These certificates eliminate unnecessary full-head scoring, and sparse writeback further reduces memory traffic and selection work. We establish the conditions for exact selection and analyze how score geometry and selection margins determine pruning effectiveness. Across English, Chinese, and code workloads, our GPU implementation accelerates the complete Indexer by up to over SGLang and over LiteTopK. End-to-end evaluation on GLM-5.2 and DeepSeek-V3.2 achieves up to and time-to-first-token speedups over the two baselines, respectively. By extending conditional computation to token selection, DoSA takes a step toward scalable attention.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.