acceptodds
Under review as a conference paper at ICLR 2027

CalibKV: Calibrated Budgets for Long-Context Attention

Abstract

Sparse attention must decide which cached keys to use and how many to retain. Fixed budgets cannot adapt to each request, while sound score bounds in our measurements admit much of the cache. We propose CalibKV, which estimates a per-head retention budget from the request's keys and selects keys by exact score at decode. A learned query subspace and a separate calibration bank set the budget during prefill; the current query determines the selected keys. We establish pairwise residual coverage under exchangeability and evaluate quality and cost at the resulting budgets. On book-passage language modeling tasks using three War and Peace segments, mean relative to Dense ranges from to over 8K to 128K contexts, while the mean budget falls from 31.2% to 2.1%. LongBench macro scores are within points of Dense on Llama and Qwen. On MATH-500, CalibKV chooses a 99.87% budget and scores 46.6%, compared with 44.6% for Dense. At long contexts, its single-layer attention calculation is faster than Dense.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.