acceptodds
Under review as a conference paper at ICLR 2027

KVShareArena: KV-Cache Reuse Across Contexts and Model Checkpoints

Abstract

Reusing key–value (KV) caches speeds up LLM inference by avoiding repeated computation on shared text. The standard method, prefix caching, reuses a KV cache only when the LLM is the same and all text before the reused content is identical. However, real workloads often break both conditions. The preceding text changes when a RAG system places different documents before the same document, or when agents with different system prompts read the same file or tool output. The LLM changes when multi-agent workflows use specialized LLMs on shared material, or when an updated model reads documents cached by its previous version. Because KV caches depend on both the preceding text and the model weights, direct reuse in these settings can reduce answer quality. Many methods repair the reused cache or compress it, each with its own balance between quality and cost, so choosing among them requires a comparison under the same conditions. However, each method paper uses its own tasks, models, and cost measures, so the reported results cannot be compared directly. Existing benchmarks do not provide such a comparison either, because they mainly test long-context processing or reuse of an unchanged prefix. To fill this gap, we introduce KVShareArena, a benchmark and open evaluation framework that compares these methods under the same conditions. KVShareArena has (1) reuse tests on 2,150 questions from three QA datasets, in which the preceding text, the LLM that wrote the cache, or both change while the answering LLM and the input text stay fixed; (2) five dense and mixture-of-experts LLMs from 4B to 30B parameters and six LLM pairs in which one version of an LLM reads caches written by another, for 33 model–dataset settings in total; (3) 11 repair and compression methods from six method classes; (4) four evaluation perspectives: answer quality, prefill computation, KV-cache memory, and latency; and (5) a common interface for adding new methods and an interactive leaderboard that lets users set their own balance between quality and speed. Experiments on KVShareArena yield two findings. First, both the quality loss from reuse and which repairs help depend on the LLM, even between two 8B models. Second, most repairs keep their quality when another version of the LLM wrote the cache, but a trained repair adapter loses quality in 12 of the 18 pair–dataset tests.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.