One Compressor Does Not Fit All
Abstract
The key-value (KV) cache dominates the memory cost of long-context inference, and many methods compress it. They are usually compared on one metric: how much task accuracy survives at a given compression ratio. Because compression is done for system efficiency, this comparison is underdetermined. Holding the model, GPU, prompts and execution stack fixed, we measure nine KV-cache configurations jointly on task score, GPU memory and latency across three benchmarks and two model scales. No configuration dominates, and the disagreement reaches the memory axis itself: attention-based eviction cuts decode-phase memory by 25% while raising whole-inference peak memory by 4.5%. Every LongBench task admits five to seven Pareto-optimal configurations, and a hindsight per-prompt oracle raises balanced utility from 0.668 to 0.714, with 87.7% of this gap left below task type. We then measure what a selector can capture from the prompt alone. Cross-fitted estimators show that the available request descriptions saturate at task granularity and that most per-prompt preference is tied to the executing model, while about 60% of the headroom concentrates in a detectable top decile. KVRoute is built on that structure: a detector over 46 tokenizer-derived features decides whether to route, a guarded ranker decides how, and refusal returns the best fixed compressor. Across three pre-specified transfer cohorts, it improves evaluation utility over FullKV by 2.82 points on average, with up to 38.2% less peak memory. This mean is dominated by ultra-long-context LongBench-v2, whose absolute 3B accuracy is below chance and which we therefore interpret as a resource-stress cohort rather than evidence of model competence. KVRoute limits severe degradation to 1.6% of requests in its worst cohort, below every fixed compressor (4.5-72.4%), and decides in 14-30ms on RULER and LongBench. Under this FullKV-relative utility and risk definition, dynamic selection combines compression efficiency with a low severe-loss rate.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.