acceptodds
Under review as a conference paper at ICLR 2027

Training-Free Latent Memory Transfer Across Heterogeneous Language Models

Abstract

Multi-model LLM systems share instructions, evidence, and interaction history across model families, yet each model keeps that context in a private KV cache. Handing it over as text forces the target to prefill the shared context again, repeating computation on every handoff. Direct cache transfer is hard: models differ in tokenization, positions, layer and head layout, and channel bases, and the target’s attention weights each cache error in a query-dependent way. A transfer must preserve the target’s computation, not merely match cache values. We present TACIT (Target-Attention-Guided Closed-Form Inter-Model Transfer), the first training-free method for KV-cache transfer across heterogeneous model architectures. TACIT aligns source and target tokens and positions, uses target attention sensitivity to select source layers, and fits per-head K/V maps in closed form. One frozen mapper is fitted per source–target direction from paired unlabeled prefills, then reused across requests and tasks so the target continues from transferred memory. Across five model families, 13 cross-family directions, and 560M–32B models, TACIT retains 83.33% of target-native accuracy on six multiple-choice tasks and outperforms all evaluated transfer baselines on every task. Cache handoff averages 6.13 ms, 6.77× faster than these baselines. On 1K–32K prefixes, end-to-end time-to-first-token improves by 1.52×–10.48× while retaining 92.4–95.4% of next-token accuracy. Analysis shows why: attention-guided selection preserves the target’s key routing, and sensitivity-weighted fitting preserves its value readout, turning model-private KV state into a reusable latent-memory interface between heterogeneous models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.