acceptodds
Under review as a conference paper at ICLR 2027

SelKV: Selective KV Cache Merging with Per-Token Merge-or-Drop and Attention Compensation

Abstract

Large Language Models (LLMs) rely on a key-value (KV) cache whose memory footprint grows linearly with context length, creating a major inference bottleneck. Token merging methods reduce this cost by consolidating evicted tokens into retained cache entries rather than discarding them, but existing approaches merge every evicted token uniformly, regardless of whether its content is compatible with its merge target, and correct the resulting softmax under-weighting of merged positions, when they correct it at all, using raw merge counts that become unstable under aggressive compression. We propose SelKV, a training-free framework with two components: a soft cosine gate that modulates merge intensity per token based on value-vector similarity, continuously interpolating between full merge and drop, and an attention-ratio compensation mechanism that corrects softmax imbalance using a bounded, prefill-derived logit bias in place of unstable count-based correction. Across three models spanning MHA and GQA architectures on 16 LongBench datasets, both components are shown to be statistically significant, and SelKV substantially outperforms existing merging baselines, with the margin widening further under aggressive compression. Against strong eviction baselines, SelKV matches performance overall while exceeding full-cache quality on complex multi-document QA, and on the RULER benchmark it retains 95% of full-cache quality at contexts up to 64k tokens. At 100k tokens, SelKV delivers a 3.3x decoding speedup over the full cache.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.