acceptodds
Under review as a conference paper at ICLR 2027

CaliTopP: Natively Trainable Top- Sparse Attention via Set-Mass Calibration

Abstract

As context windows grow, the quadratic cost of full attention becomes the main bottleneck of long-context LLM inference, motivating sparse attention. Learning lightweight indexers for token selection is emerging as a mainstream approach because it maintains consistency between training and inference. However, these indexers almost universally use fixed top-k budgets, overlooking the vast differences in token requirements across tasks and models. In contrast, top-p selection retains a target fraction of attention mass, providing a more reliable accuracy guarantee and an input-adaptive budget. Top-k selection depends only on token rankings, whereas top-p additionally depends on cumulative probability mass. We show theoretically and empirically that KL-distilled top-k indexers can accurately rank tokens yet underestimate the cumulative mass of important tokens, leading to severe over-selection when reused for top-p selection. Based on this finding, we propose CaliTopP, a trainable token-level top-p sparse-attention framework. Its parameter-free budget-calibration loss aligns the indexer's selected budget with the model's true requirement without adding parameters or reducing the recall of important tokens. With only two lightweight training stages, CaliTopP transfers across diverse architectures and achieves up to a speedup over dense attention while preserving generation quality.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.