acceptodds
Under review as a conference paper at ICLR 2027

Where Does Attention Mass Live? A Matched-Budget Study of Query-Adaptive Attention Sparsification for Decoder LMs

Abstract

Attention sparsification is the dominant route to cheap long-context inference, yet practitioners choosing a masking rule face a crowded menu of heuristics with little guidance on which buys the most quality at a fixed kept-fraction budget. We present a controlled, matched-budget comparison of seven sparsification families—query-adaptive thresholding in logit and softmax-probability space, per-query top-k, sliding window, sink-plus-window, static block-sparse, and random masking—evaluated by exact masked-softmax teacher-forced loss on Qwen2.5 models at 0.5B and 1.5B parameters, across WikiText-2 and C4, at context lengths from 1,024 to 8,192. Three findings emerge. First, thresholding in probability space (retain entries whose softmax weight exceeds τ times the per-query maximum) matches full-attention loss within 0.007 nats/token at a realized kept-entry fraction of 13.5% at context 1,024 and within 0.030 nats at 4.1% at context 8,192 (fp32 paired residual: 0.008, CI [0.005, 0.010]); in the budget-matched comparison (8K context, n=8 paired documents, 4.1% versus 4.6% kept fraction) it beats magnitude-based top-k by 0.026 nats (paired 95% CI [0.018, 0.035]) at the lower budget. Second, the same rule defined in logit space—the natural first proposal—is consistently dominated, an honest negative result we trace to the multiplicative logit-ratio rule interpolating toward the row minimum rather than a calibrated gap from the maximum. Third, per-head budget profiles measured on one set of documents transfer almost perfectly to unseen documents (mean pairwise Pearson r = 0.98), but profile-guided reallocation—per-layer or per-head— does not improve on the uniform threshold (it costs +0.028 nats, significant on every paired test): the sparsity pattern is a stable property of the model, and a uniform threshold is already near-optimal in this family under proportional reallocation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.