acceptodds
Under review as a conference paper at ICLR 2027

Can Accurate Prediction Tell Sparse Attention What to Keep in LLMs?

Abstract

Sparse attention reduces long-document cost by passing only a few blocks to later computation. We ask whether accurate next-token prediction from one source identifies what a router should retain under distribution shift. In general, it cannot because labels from one source and unlabeled target documents can leave several deployment worlds compatible with the observations. Repeated facts and visible route positions also prevent a unique mask from being the right target. We characterize the resulting possible and necessary selector sets under a fixed retention budget and declared source, selector, predictor, and loss classes. Labeled source variation recovers the optimal set when suboptimal rules are uniformly separated in worst-source risk, and source mixtures give exact population certificates under convex predictor classes. Label-preserving masking gives a second route with multi-view and adaptive guarantees, whose cost is controlled by informative-example rates and squared performance gaps. Matching lower bounds clarify what extra data are needed before pruning long context and why predictive success alone does not establish causal evidence or computational savings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.