acceptodds
Under review as a conference paper at ICLR 2027

Which Key-Value Cache Compressions Can Continued Pretraining Repair?

Abstract

Retrofitting pretrained transformers with smaller key–value (KV) caches can reduce memory use and accelerate decoding with limited additional training. Next-token prediction incentivizes the use of context, but whether continued pretraining restores retrieval after cache compression is unclear. We compare head merging, joint low-rank compression of keys and values, and token pooling at matched cache sizes and training budgets on Qwen3 from 0.6B to 8B and on Falcon3-1B, and extend the merge to five model fameters. Each conversion isevaluated against an unconverted control trained with the same recipe. Recovery in language-modeling loss and commonsense performance does not reliably predict multi-key retrieval, whichrget key from many similarentries. At 4K tokens and half the cache, a low-rank latent retains 97.9% of the control's multi-key retrieval at 4B, but head merging only 34.0%, despite similar commonsense performance. Ahe retrieval headsidentified in one forward pass also improves recovery at the training length, at a smaller cache saving, but the benefit is gone at 16K tokens. Continued pretraining also fails on hing. Rotating value headswithout adjusting their output projections permits an exact inverse correction, yet continued pretraining lowers the loss without restoring retrieval. From 0.6B to 8B, a closed-form localst 93.6% of the control'smulti-key retrieval without continued pretraining. These results motivate three principles for KV-cache retrofitting: evaluate multi-key retrieval explicitly, preserve retrieval-relevann misalignments locallybefore continued pretraining.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.