acceptodds
Under review as a conference paper at ICLR 2027

Know Before You Transfer: Predicting and Certifying Cross-Model KV Cache Reuse

Abstract

Multi-model inference pipelines repeatedly process the same context when requests move between models. At each handoff, the receiving model typically prefills the shared context from scratch, even though the sender has already processed it. Existing methods reduce this cost by transferring the sender's key-value (KV) cache into the receiver's cache space through a learned or closed-form map. Their effectiveness, however, varies widely across model pairs, and failures are hard to anticipate or detect: a transferred cache may yield a fluent answer that disagrees with the output of a full prefill. Reliable deployment therefore needs inexpensive answers to two questions: which model pairs admit effective transfer, and how much risk a transferred cache poses for a given request. We address both through a first-order analysis of how cache perturbations propagate through attention. The analysis shows that downstream deviation depends on where the reconstruction error lies relative to the receiver's query geometry, a dependence that error magnitude alone does not capture. From it we derive the **Cache Transferability Index**, an offline diagnostic that predicts a pair's achievable retention from second-moment statistics alone, without fitting a mapper. We also develop **CacheCert**, an online procedure that turns the same geometry into a per-request risk score and applies conformal risk control to obtain a distribution-free bound on the rate of unsafe transfers under exchangeability. Across eighteen directed model pairs from three families and five long-context workloads, the index reaches a Spearman correlation of 0.91 with retention at roughly one fortieth the cost of fitting and evaluating a mapper. CacheCert keeps empirical risk at or below its nominal target on every pair. At a 10 percent risk budget, it saves 66 percent of prefill work on LongBench multi-hop QA and reduces silent failures from 14.7 percent to 4.1 percent on an agent-report track. Furthermore, both components are backbone-agnostic and can compose with existing transfer and repair methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.