acceptodds
Under review as a conference paper at ICLR 2027

Double-P: Hierarchical Top-P Sparse Attention for Long-Context LLMs

Abstract

Attention over a growing key-value (KV) cache dominates the cost of long-context decoding. Top- sparse attention keeps the smallest set of tokens whose attention mass reaches a target , so the budget adapts to each head and decoding step. Existing top- methods, however, use top- only to prune: a fixed-budget selector first proposes candidate tokens, and token-level score estimation then removes some of them. Such pipelines cannot recover attention mass that the candidate set misses, and their estimation cost grows with context length. We present Double-P, which uses top- to select. Double-P estimates the attention mass of every KV cluster from size-weighted centroids and applies top- over the entire context, so the attended set is sized by the attention distribution rather than by a budget. A second top- threshold on the same estimates allocates precision: high-mass clusters receive exact token-level attention, and the remaining selected clusters are approximated by their centroids instead of being discarded. Across long-context benchmarks, Double-P consistently achieves near-zero accuracy drop, reducing attention computation overhead by up to 1.8 and delivers up to 1.3 end-to-end decoding speedup over state-of-the-art fixed-budget sparse attention methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.