acceptodds
Under review as a conference paper at ICLR 2027

SCALE-KV: State-Augmented Cross-Scale KV Cache Reuse for Long-Context Inference

Abstract

Long-context inference and context handoff in multi-agent collaboration incur substantial context-processing costs. Cross-scale key-value (KV) cache reuse allows a large model to reuse context computation performed by a smaller model, reducing repeated large-model prefilling and redundant computation after model handoff. However, structural and representational differences across model scales can degrade generation quality, while additional computation for quality recovery may offset the resulting speedup. We propose SCALE-KV, a state-augmented framework for cross-scale KV cache reuse. SCALE-KV establishes cross-model correspondence through a fixed base mapping and learns compact state-augmented residuals from the source KV cache and intermediate states. The cache converter and low-rank target-reader adaptation are jointly optimized for generation while both backbones remain frozen. At inference time, the source model prefills the full context once, after which the target directly processes questions and generates answers from the converted prefix without repeating full-context prefilling. Across three question-answering evaluations with 8–16K contexts, SCALE-KV remains within 3.53 F1 points of native large-model inference. Including source-model prefilling, SCALE-KV achieves up to time-to-first-token (TTFT) speedup for fresh requests; when the source cache and required intermediate states are available, it achieves up to TTFT speedup for handoffs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.