Read Once, Reuse Across LLM Families: KV Cache Sharing with CachePort
Abstract
Multi-model LLM applications repeatedly process the same document or conversation when passing context between models. Reusing key–value (KV) caches could avoid this work, but differences in tokenization, layer and head structure, and learned representations prevent direct reuse across model families. We introduce **CachePort**, a framework for cross-family cache reuse that avoids constructing a complete native receiver cache. CachePort aligns states at shared text boundaries, translates them through measured layer correspondences and full-width affine maps, and combines a lightweight receiver adapter with limited native recomputation. Both models' pretrained weights remain frozen; only the translators and the adapter are trained. Experiments spanning Llama, Qwen, and Mistral demonstrate useful context transfer for language modeling, question answering, and summarization, with quality depending on the transfer direction and recomputation budget. A receiver adapter also supports a source excluded from its training when paired with newly fitted translators. For Llama-3.1-8B to Mistral-7B transfer, an optimized receiver-computation path achieves speedups over full prefill across tested context lengths, excluding serving-system overhead. Separately, a vLLM integration reduces receiver handoff latency from 123.9 to 75.8 ms at an 8k-token context with a perplexity increase of 0.30 relative to full prefill. These results establish a practical quality–latency trade-off for cross-family context reuse.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.