acceptodds
Under review as a conference paper at ICLR 2027

Freeze the Decoder, Heal the Encoder: Parameter-Efficient Adaptation for SVD-Based KV-Cache Compression

Abstract

The KV cache dominates inference memory, especially for vision-language models, which must cache image and text tokens alike. A pretrained model can be retrofitted with a compact, multi-head-latent-attention-style cache by SVD-factorizing its key/value projections into a down-projection ("encoder") and an up-projection ("decoder"), followed by a short fine-tune ("healing") that recovers the accuracy lost to truncation. Existing conversions heal both factors, without asking whether both need it. We freeze the decoder at its SVD initialization and heal only the encoder, which trains 1/3 of the new parameters. This keeps keys and values in the original top-r output subspace, but leaves the model free to choose what the cache stores. Under per-arm learning-rate tuning and three seeds, encoder-only healing is competitive with decoder-only and full healing on Qwen2.5-VL-3B-Instruct (MME, TextVQA) and on average accuracy over six zero-shot text tasks across two 7B backbones, including a setting with 3.5-6.3 points of headroom left. No difference is statistically significant, and equivalence tests bound any difference to within 1.2-5.4% of the reference score per metric; with three seeds, smaller deficits cannot be ruled out. Encoder-only healing uses 3x fewer trainable parameters and 3x less optimizer-state memory (0.57 vs. 1.72 GiB), producing an identical served model. We also find that comparing healing modes at one shared learning rate makes encoder-only healing look far better than it is, with gaps that pass a 3-seed t-test (t up to 8.78) but vanish under per-arm tuning, a pitfall for comparing any fine-tuning recipes of unequal size.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.