PANORAMA: CROSS-MODEL KV SHARING VIA CONTEXTUAL RECOMPUTATION
Abstract
Real-world LLM workloads, such as multi-agent systems and model cascades that escalate requests to larger models, frequently switch a request between models while retaining the same context. Because KV caches are tied to model parame- ters and cannot be shared across models, each switch forces the target model to re-prefill the entire context that the source model has already processed, incur- ring substantial computation cost and increasingly delaying time to first token as context length grows. Existing cross-model KV translation methods avoid this re-prefill by mapping the source KV cache into an approximate target KV cache, but translation alone can substantially degrade accuracy. Selective KV recomputa- tion has been shown to repair context-inconsistent KV cache entries produced by the same model, but whether selective recomputation remains effective for a KV cache translated from a different model is unknown. Our key observation is that even when a translated KV cache is insufficient for direct decoding, the translated cache can still provide useful context for target-model recomputation. Based on this observation, we propose PANORAMA, a cross-model KV sharing method that recovers high-quality target inference from approximate cross-model KV with- out full target re-prefill. PANORAMA contextually recomputes a small subset of important target positions over the full KV cache produced by the cross-model mapper, allowing target-model computation to exploit the translated context. To provide the full translated context efficiently, PANORAMA factorizes the mapper into a shared source projection and layer-specific target reconstruction, transfer- ring only a compact latent per prompt position. PANORAMA further refines the mapper under the repaired mixed cache used at inference, removing the training– inference discrepancy between translation-only mapper fitting and mixed-cache decoding. Across model pairs and tasks, PANORAMA achieves the highest ac- curacy among prior cross-model handoff methods on nearly all benchmarks at comparable handoff cost. Moreover, in the Qwen3 8B-to-32B switching scenario, our vLLM implementation of PANORAMA reduces handoff latency by 5.84×over native re-prefill at 16K tokens with minimal accuracy loss.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.