Cross-Model KV Cache Construction in LLM Familiies: Residual-Stream Mapping Without Matched Head Geometry
Abstract
Passing context between language models usually means serialising it as text and prefilling it again. We present the residual-stream map (RSM), a method for build- ing one model’s key-value (KV) cache from another model’s internal states that is head-geometry-independent within tokenizer-compatible model families. A per- layer ridge map carries the sender’s pre-normalisation residual states into the re- ceiver’s residual space, a small corrector trained on generic text adjusts them, and the receiver’s own frozen normalisation, key/value projections and rotary encod- ing build the cache, so its head count and head dimension are always the receiver’s. We evaluate every directed pair among finetunes and sizes of Qwen and Llama models from 0.6B to 8B parameters. Transfer between Llama-3.2-1B (head di- mension 64) and Llama-3.2-3B or Llama-3.1-8B (head dimension 128), which KV-space mappers cannot express, runs unchanged: into the 1B model it scores 39.36 and 38.02 F1 on HotpotQA against 40.28 for its own text prefill (97.4% and 93.5% retention). Between finetunes of one model, transfer retains a median 102.5% (Llama) and 90.4% (Qwen2.5) of the F1 gain of text prefill, and across sizes it reaches or comes within 7 F1 of the weaker model’s own text reading in 10 of 12 directions. Building the cache is 3.1 to 5.3 times faster than prefilling 1024 tokens. On the pairs where training-free KV-transfer methods also run, RSM matches them between finetunes (40.47 against 42.41 F1 for the best mapper on Llama) and leads them by 10.31 to 31.34 F1 across sizes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.