acceptodds
Under review as a conference paper at ICLR 2027

RankKV: NDCG-Bounded KV Reranking against Attention Foam

Abstract

Long-context serving must evict KV positions once the cache exceeds a budget, and the default keep rule ranks tokens by accumulated attention. We find that attention greed does not degrade every prompt uniformly under a tight cache. A few structurally prominent tokens monopolize the keep set, and quality versus Full KV—the uncompressed cache—is bipolar: some prompts stay as good as Full KV, or even exceed it, while others collapse far below. We call this allocation attention foam and trace the Matthew split to single-surrogate eviction: most compressors still rank by one score, so they retune monopoly rather than leave it; the keep set is a local optimum of that score. However, replacing attention by another single criterion is the other pole, not the remedy. An effective correction must therefore rewrite the host compressor's greedy keep set at controllable strength, without abandoning that compressor. In this paper, RankKV is a training-free post-process behind the host: it swaps a bounded number of high-value evicted tokens into the greedy keep set while preserving a floor on greedy attention coverage. Behind , PyramidKV, SnapKV and Ada-SnapKV, across three model families and three long-form tasks, the method is net-positive per model. Mean lift is redistribution: healthy prompts stay near the host; hurt prompts recover.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.