acceptodds
Under review as a conference paper at ICLR 2027

Efficient Attention via Pre-Scoring: Prioritizing Informative Keys in Transformers

Abstract

Efficient attention methods restrict which key-query pairs are evaluated, but the rules they use are query-dependent and data-independent: HyperAttention hashes queries and keys with angular locality-sensitive hashing, then estimates everything outside the resulting blocks by sampling residual columns uniformly. We replace that uniform proposal with a query-independent, data-dependent prior over keys, ranked once per layer by clustering or leverage-style scoring. We bound the resulting output error as a tail-mass bias plus a variance term linear in the retained budget, which predicts the interior optimum we observe. On ChatGLM3-6B-32k with LongBench, pre-scoring reduces perplexity from to at a retained budget of keys; clustering dominates leverage-score selection at tight budgets ( vs. at keys) and converges to it as the budget grows. In ViT-Large it retains \verb|98.4%| of base accuracy at keys per head. Theoretically, we show that the importance measure matched to degree- polynomial attention is the sensitivity-and that Minkowski- clustering recovers the top--sensitivity keys under a planted-subspace model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.