CaliTopP: Natively Trainable Top- Sparse Attention via Set-Mass Calibration
Abstract
As context windows grow, the quadratic cost of full attention becomes the main bottleneck of long-context LLM inference, motivating sparse attention. Learning lightweight indexers for token selection is emerging as a mainstream approach because it maintains consistency between training and inference. However, these indexers almost universally use fixed top-k budgets, overlooking the vast differences in token requirements across tasks and models. In contrast, top-p selection retains a target fraction of attention mass, providing a more reliable accuracy guarantee and an input-adaptive budget. Top-k selection depends only on token rankings, whereas top-p additionally depends on cumulative probability mass. We show theoretically and empirically that KL-distilled top-k indexers can accurately rank tokens yet underestimate the cumulative mass of important tokens, leading to severe over-selection when reused for top-p selection. Based on this finding, we propose CaliTopP, a trainable token-level top-p sparse-attention framework. Its parameter-free budget-calibration loss aligns the indexer's selected budget with the model's true requirement without adding parameters or reducing the recall of important tokens. With only two lightweight training stages, CaliTopP transfers across diverse architectures and achieves up to a speedup over dense attention while preserving generation quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.