acceptodds
Under review as a conference paper at ICLR 2027

LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models

Abstract

A small language model may have processed thousands of tokens before a router calls a larger model. Transferring attention key–value (KV) caches can avoid repeated prefix processing, but hybrid models also maintain recurrent memory. Can a differently sized receiver use this persistent state? We study Qwen3.5 4B→9B handoff without historical target replay. Adding the transferred Gated DeltaNet persistent-state package to fixed translated KV lowers teacher-forced continuation loss by 0.747 nats/token (95% paired document bootstrap CI [0.692, 0.805]), improving all 64 held-out books. Direct recurrent/convolution reuse outperforms the tested learned maps despite lower reconstruction error for the learned recurrent map. Without learned correction or target repair, a separate replication finds native context-benefit recovery of 0.833 at 4K and 0.844 at 16K using post-trained Q4_K_M models in llama.cpp, compared with 0.142 and 0.105 for KV-only transfer, across four domains. A compact, pair-specific residual correction separately reduces Base-model excess loss to 0.076 nats/token, recovering 0.918 of the native target's context benefit while beating continued 4B prediction. Small trajectory perturbations also expose transfer-specific degradation. These results establish useful but conditional portability of built-in persistent state within one geometry-matched model family. They do not establish native-equivalent free generation, cross-family compatibility, downstream-task equivalence, or production speedup.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.