Count Competitors, Not Context: When KV-Cache Budgets Transfer
Abstract
Every KV-cache eviction method must decide how many tokens to keep, and practice sets that budget as a fraction of the prompt. We show that the fraction is the wrong unit. The budget is set by competition: evidence survives eviction exactly when the budget covers the evidence plus the tokens the scorer ranks above it, and prompt length matters only through those competitors. Prior work misses this for two reasons. It evaluates budgets on prompts where length and competing content grow together, so their effects cannot be separated, and it treats keeping the evidence as enough to answer. Separating length from competition exposes costly errors in both directions: on long prompts with little competing content, a fixed fraction keeps about 100 times more cache than needed, while at a fixed length, added distractors raise the required budget more than 200-fold. Counts frozen before verification forecast budgets within a factor of two at up to 512K tokens, where a constant fraction fails. Keeping the evidence is also not enough: answers depend on keeping the rows that bind the evidence to the question. Compacting the cache to these budgets cuts KV memory by up to two orders of magnitude with no loss of accuracy on the tested requests. The rule for sizing a cache is simple: count competitors, not context.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.