acceptodds
Under review as a conference paper at ICLR 2027

The Compression Square: Weight Quantization and KV-Cache Eviction Do Not Compose at the Operating Point

Abstract

A deployed long-context language model is compressed twice: its weights are post-training quantized, and its KV cache is pruned by an eviction policy that ranks tokens with the model's own attention. Each axis is evaluated alone and the two losses are assumed to add, which misses that the scores deciding what to evict come from the quantized model. We measure the coupling with the compression square, a factorial over model precision and selection source whose counterfactual transplant cell splits the interaction exactly into an oracle term (the quantized model kept the wrong tokens for itself) and a fragility term (it is more hurt by the same eviction). Across 499k released per-item records on Qwen2.5-7B, Llama-3.1-8B and Qwen2.5-3B at 4k–32k, the deployed cell is significantly worse than additive in 71 of 143 identifiable SnapKV cells (63 after Benjamini–Hochberg) and sub-additive in 3; on the needle tasks the effect is a tail: most items are unaffected and a few are badly hurt. Its size is set by the quantizer (two noisy RTN builds exceed the three shipped W4-g128 recipes in 32 of 33 matched conditions), but the shipped recipes are not exempt on real documents: on Llama-3.1-8B LongBench, GPTQ W4-g128 is super-additive in 9 of 9 cells over three eviction policies at 288 items (5 at 120), and on Qwen2.5-7B LongBench, where eviction alone costs 0.7% of the baseline loss and RTN W4-g128 +0.07 nats, the two together cost +0.49 [+0.14, +0.92] more than their sum. An intervention supports the channel: i.i.d. noise injected into the ranked scores alone, with the model untouched, opens it at a level within a factor of two of a quantizer's own score error. At these budgets, single-axis evaluation does not predict the deployed system. All code and the 499,289 per-item records will be made public upon acceptance at https://github.com/[anonymized]/[anonymized].

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.