acceptodds
Under review as a conference paper at ICLR 2027

Read Once, Reuse Across LLM Families: KV Cache Sharing with CachePort

Abstract

Multi-model LLM applications repeatedly process the same document or conversation when passing context between models. Reusing key–value (KV) caches could avoid this work, but differences in tokenization, layer and head structure, and learned representations prevent direct reuse across model families. We introduce **CachePort**, a framework for cross-family cache reuse that avoids constructing a complete native receiver cache. CachePort aligns states at shared text boundaries, translates them through measured layer correspondences and full-width affine maps, and combines a lightweight receiver adapter with limited native recomputation. Both models' pretrained weights remain frozen; only the translators and the adapter are trained. Experiments spanning Llama, Qwen, and Mistral demonstrate useful context transfer for language modeling, question answering, and summarization, with quality depending on the transfer direction and recomputation budget. A receiver adapter also supports a source excluded from its training when paired with newly fitted translators. For Llama-3.1-8B to Mistral-7B transfer, an optimized receiver-computation path achieves speedups over full prefill across tested context lengths, excluding serving-system overhead. Separately, a vLLM integration reduces receiver handoff latency from 123.9 to 75.8 ms at an 8k-token context with a perplexity increase of 0.30 relative to full prefill. These results establish a practical quality–latency trade-off for cross-family context reuse.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.