acceptodds
Under review as a conference paper at ICLR 2027

Compression-Aware Fine-Tuning: Adapting Language Model Weights to a Frozen Cache Bottleneck

Abstract

The low-rank projection of the key-value (KV) cache is a natural strategy to reduce inference memory. However, compressing a model whose weights were never trained to tolerate the resulting information loss leads to severe quality collapse at moderate compression ratios. Compression-aware fine-tuning (CAFT) closes this gap by adapting the model weights to a frozen cache compressor rather than compressing weights that were never trained for the loss. At a 2x cache budget on Llama-3.2-1B, CAFT reduces WikiText-103 perplexity by 11.5x relative to the strongest post-hoc baseline (a fine-tuned checkpoint compressed under a basis freshly calibrated on its own activations). The advantage holds across six compression operators, three model scales (1B to 8B), and two training domains, and persists when stacked with 4-bit code quantization. At 2x and 3x compression, CAFT preserves single-needle retrieval and maintains a large advantage on LongBench where the post-hoc baseline collapses. Mechanistically, the gain is functional robustness: the network routes around the discarded dimensions rather than concentrating cache energy into the retained subspace, as shown by spectral probes and basis-transfer controls that falsify induced compressibility. Through this paper, we demonstrate that CAFT recovers the capability that low-rank KV cache compression destroys.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.