acceptodds
Under review as a conference paper at ICLR 2027

SIEVE: Estimate-Aware Clustering in the Query-Induced Space for Sparse Attention

Abstract

Sparse attention uses only a small part of the KV cache at each decoding step, while cluster-based indices rank and estimate groups of keys through their centroids. Existing methods cluster keys by angular similarity, even though attention depends on the inner products between keys and decoding queries. Keys that are close in angle can therefore induce very different logits, causing important keys to be diluted in both retrieval and estimation. We introduce **SIEVE**, which clusters keys in the score space induced by the query distribution. Its objective minimizes residual logit error under a robust query-induced metric and places additional weight on keys that are likely to become important retrieval targets. We show that this objective controls errors in attention estimation. Across long-context and long-reasoning benchmarks, SIEVE outperforms training-free sparse-attention baselines at equal budget and remains close to full attention while using only a small fraction of the KV cache.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.