acceptodds
Under review as a conference paper at ICLR 2027

Ranking Is Not Support Selection: Exact Marginals and Certified Headroom for Sparse Attention in LLMs

Abstract

Long-context large language model (LLM) inference often reduces attention memory or bandwidth by retaining or fetching only a limited subset of cached tokens. A common abstraction is to score candidates independently, keep the Top-, and renormalize attention over the retained set. We ask a basic question: if those per-token scores were exact, would Top- recover the best -token support? In general, it does not. Even when every candidate is scored by its exact downstream singleton-deletion Kullback–Leibler (KL) divergence under the same next-token objective used to evaluate the retained set, an explicit Top-2 example incurs the KL of the optimal pair. The reason is structural: softmax renormalization makes the value of adding or replacing a token depend on which other tokens are retained. We derive the exact exchange identity and prove support-conditioned preference reversals for every . We then test this distinction by exhaustive support enumeration in softmax-attention heads of SmolLM2, Gemma-2-2B, and Qwen3.5-4B. At the primary budget, exact-singleton Top- leaves at least 10% relative headroom to the exact optimum within the same candidate universe in 55/60, 49/60, and 60/60 prompt–head cells, respectively. The effect persists across alternative candidate-universe constructions, while its magnitude varies substantially across models. Direct mechanism analysis observes the predicted support-conditioned reversals in nearly every substantial-headroom case. Finally, \SwapCert searches the same-budget one-swap neighborhood for a concrete witness of remaining headroom and detects 52/55, 48/49, and 57/60 exact positives at . These results separate two questions that can be conflated in sparse LLM inference: which tokens are important individually, and which fixed-budget set best preserves the downstream next-token distribution.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.