acceptodds
Under review as a conference paper at ICLR 2027

CACHET: How to Make KV-Cache Sharing Carry What the Larger Model Read

Abstract

Collaborative inference lets a small language model use the understanding of a larger one. Rather than passing that understanding through lossy text, recent methods share the larger model's KV cache, reshaped by a learned translator for the small model. When both models read the same prompt, accuracy rises, as if the small model had received what the larger one read. Yet the state-of-the-art C2C fuser still retains half of its gain when handed the cache of an unrelated question, and its updates track the question rather than the cache. The gain is largely an effect of added parameters adapting the small model's own cache, which the fuser also reads. The collaboration is then capped at what an adapter can add, however much the larger model understood. We propose CACHET, a translator designed so that the small model's answer depends on what the larger model read. It reads only the larger model's cache, closing that shortcut, and is trained to reproduce, jointly across all heads, keys, and values, the cache the small model would have built from the same input. Because most of the key cache is common to all inputs, this training emphasizes the part that differs between inputs. The loss is kept throughout answer training, which alone would pull the transmitted cache off target. When only the larger model sees the passage that determines the answer, a 0.6B model assisted by a 4B one reaches 97.3% without training on this test and drops to 1.0% when given the cache of a passage supporting a different answer. On four benchmarks, CACHET recovers 87.5% of the accuracy gap to the larger model, against 36.0% for the C2C fuser.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.