CacheMorph: Asymmetric KV Transfer for Efficient Model Switching
Abstract
Multi-model LLM systems often move a request between models, and routers switch models as load changes. Each switch has a hidden cost: the receiving model must re-prefill a shared context that may be tens of thousands of tokens long. Reusing the source model's key-value (KV) cache could remove this cost. However, same-model reuse methods do not directly support such switches, while cross-model cache alignment can still incur substantial accuracy losses. Aligning KV representations alone does not ensure accurate receiver outputs. We show why this gap matters: key errors change where attention reads, value errors change the content it retrieves, and query heads sharing KV respond differently to the same errors. Based on these insights, we introduce CacheMorph, a framework that enables the receiver to skip prefill of shared history by translating KV caches across models. Its asymmetric mapper predicts keys and values separately, conditioning each value on its mapped key. A lightweight, query-adaptive reader then adjusts how each head accesses the shared cache. Both components are trained to match the frozen receiver's native-cache output distribution. A streaming pipeline overlaps cache transfer and mapping with ongoing computation. CacheMorph improves accuracy over prior cross-model KV reuse methods, especially across model families. At 32K tokens, it achieves up to a time-to-first-token speedup over receiver re-prefill. In RAG cascades, it delivers approximately speedup and over 50% lower estimated token cost while retaining over 90% of re-prefill accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.