What Transfers Across Models in KV Cache Eviction? A Controlled Measurement Study
Abstract
Deploying a KV cache eviction method on an untuned model is a leap of faith. Whether published accuracy holds on a new model and a new task is unknown. We separate eviction into independent design choices and measure each on five models in one physical-eviction harness. For scoring, accuracy follows a computable property of the workload rather than of the model. Viewing eviction as estimation yields a parameter-free dilution regularity. Windowed scoring attenuates the task-conditional signal by the fraction of task rows in the window. It holds on all five models under controlled interventions in both directions and on real tasks spanning question answering, summarization, and long-document understanding. For allocation, per-layer fragility follows no cross-model pattern, so borrowed shapes must be measured rather than trusted. A starvation probe maps fragility from a few calibration prompts and tests borrowed shapes against a uniform split. Derived only from controlled measurements, this knowledge predicts where the complete methods CAKE and ChunkKV fail on a new model and prescribes a repair, lifting needle-in-a-haystack retrieval from 43% to 65% with 128 tokens per layer at 32K context. What transfers across models is not a method but knowledge of what transfers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.