acceptodds
Under review as a conference paper at ICLR 2027

Consistency over Accuracy: Outlier-Bypass Layer-Adaptive Quantization for Codec-Based KV Cache Compression

Abstract

Reusing key–value (KV) caches can eliminate repeated prefill computation in long-context large language model serving, but practical reuse requires the cache to be compressed without significantly altering the model response. Practical lossy KV cache compression raises two challenges: how to measure reuse fidelity and how to mitigate quantization distortion. For the former, aggregate task accuracy is insufficient to reliably characterize answer consistency before and after compression, as compression-induced perturbations may spuriously turn an incorrect answer into a correct one. To address this issue, we propose replacing accuracy with answer consistency as the metric for evaluating compression-induced distortion. For the latter, existing codec-based systems compress the dense representation effectively, yet sparse outliers are not adequately handled. To be specific, a few extreme values can consequently dominate the shared group scale, reducing the effective resolution of normal values and compromising subsequent related allocations. We refer to this effect as outlier-scale pollution. Accordingly, we propose outlier-bypass layer-adaptive quantization. The method detects group-wise extremes by gap prominence, stores them exactly in a separate sparse bypass. Then, the remaining values are quantized using layer-specific level where the corresponding thresholds are customized for each model through offline cumulative-distribution fitting (CDF) calibration. By the calibrated CDF, we derive two presets corresponding to distinct consistency–compression operating points. Across five models and three long-context benchmarks, the high-fidelity preset (Ours-HF) attains 82.33% consistency at 7.41× compression, simultaneously outperforming KVCodec and CacheGen by +1.96 percentage points in consistency with +5.26% and +26.24% in compression ratio, respectively. For workloads that prioritize storage efficiency, the high-compression preset (Ours-HC) further increases the compression ratio to 8.35×, a gain of 18.61% over KVCodec and 42.25% over CacheGen, at a modest consistency cost of 1.45 percentage points.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.