KVRelay: Reusing KV Caches Across Language Models without Re-Prefilling
Abstract
Deployed language models reuse context far more often than they compute it: the same document, prompt or session is encoded once and read back many times, which turns its KV cache into stored infrastructure. Model iteration invalidates that infrastructure. The day a deployment upgrades, every cached context still lives in the previous model’s attention space, and the new model has to prefill all of it again. We ask whether a target model can inherit a KV cache produced by another model and continue inference without re-prefilling the context. The obstacle is that a cache is an interface rather than a portable encoding of text: even between models that share a tokenizer, depth and KV geometry, injecting the source cache directly leaves the target below its own no-context accuracy on four of five pairs. We introduce KVRelay, which treats cross-model reuse as interface adaptation. It supervises a translated cache by the computation it induces inside frozen target attention rather than by element-wise reconstruction of the target’s native KV tensors; it starts from a parameter-free, target-calibrated base map and learns a zero-initialized cross-head correction; and it gives the Value branch a short causal mixer over source history. Across five Qwen/Llama model pairs and five context-dependent benchmarks, KVRelay recovers 90.3 ± 1.0% of the tar- get’s native context gain over five seeds with a translator worth 0.3–5.1% of target parameters; replacing functional supervision with direct KV MSE reduces five- pair mean accuracy by 4.88 percentage points. Complete translation costs 13.6% of target prefill on average and reduces end-to-end time to first token by 3.41×. KVRelay turns a model upgrade from a full cache rebuild into a cache handoff.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.