acceptodds
Under review as a conference paper at ICLR 2027

Easy-Kuant: Diagnosing and Mitigating KV-Cache Quantization Collapse with Selective Layer Protection

Abstract

Serving large language models increasingly relies on quantizing the key–value (KV) cache to hardware-native low-bit formats such as NVFP4 and FP8, widely reported to be near-lossless at four bits. On the Qwen2.5 family and its derivatives the same one-line switch is catastrophic: in vLLM, Qwen2.5-7B-Instruct falls from 90.2% to 11.0% on GSM8K under NVFP4 and to 2.2% under FP8, and vision-language checkpoints fall to 0%. Rather than tuning against cache reconstruction error, as prior methods do, we ask *where* the error becomes a failure. Replaying each layer's attention from clean activations shows that cache error is flat across depth while attention-output error varies by an order of magnitude, enters almost entirely through keys, and concentrates in one to four layers; two-directional interventions show that these layers are sufficient both to cause and to repair the collapse. The diagnosis becomes Easy-Kuant: a few forward passes yield one statistic per layer, a fixed threshold turns it into the list of layers to keep in BF16, and stock vLLM applies the list through an existing argument—no kernel, no search, no gradients. On the native serving path Easy-Kuant returns NVFP4 to 90–96% of BF16 accuracy on the 3B and 7B checkpoints and FP8 to 98–100%, transfers unchanged to long-context and multimodal checkpoints, matches software baselines that cannot run in vLLM, and costs 0.94× the per-token KV processing of full NVFP4 at 85% of its cache capacity. An optional quantization-aware stage, trained on the model's own rollouts and guided by the same layer list, adds up to 3.7 further points where a gap remains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.