acceptodds
Under review as a conference paper at ICLR 2027

Count Competitors, Not Context: When KV-Cache Budgets Transfer

Abstract

Every KV-cache eviction method must decide how many tokens to keep, and practice sets that budget as a fraction of the prompt. We show that the fraction is the wrong unit. The budget is set by competition: evidence survives eviction exactly when the budget covers the evidence plus the tokens the scorer ranks above it, and prompt length matters only through those competitors. Prior work misses this for two reasons. It evaluates budgets on prompts where length and competing content grow together, so their effects cannot be separated, and it treats keeping the evidence as enough to answer. Separating length from competition exposes costly errors in both directions: on long prompts with little competing content, a fixed fraction keeps about 100 times more cache than needed, while at a fixed length, added distractors raise the required budget more than 200-fold. Counts frozen before verification forecast budgets within a factor of two at up to 512K tokens, where a constant fraction fails. Keeping the evidence is also not enough: answers depend on keeping the rows that bind the evidence to the question. Compacting the cache to these budgets cuts KV memory by up to two orders of magnitude with no loss of accuracy on the tested requests. The rule for sizing a cache is simple: count competitors, not context.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.