KVBench: An Extensible Benchmark to Evaluate Approximate KV Cache Reuse
Abstract
Approximate KV-cache reuse accelerates large language model inference by reusing cached states beyond identical request prefixes. However, these states depend on their original context, so transferring them can alter generation quality, and the resulting quality–latency trade-off may vary across models, tasks, and request structures. We present KVBench, an extensible benchmark that separates tasks, models, reuse methods, and request workflows to support controlled evaluation of these effects. KVBench measures task quality and online time-to-first-token relative to full-prefill references across RAG-style context reuse, multi-agent communication, and agentic skill reuse, covering five models with full- or hybrid-attention architectures. Our evaluation reveals substantial task–model variation in quality retention and recovery through recomputation. Controlled studies further show that instruction placement can strongly affect approximate reuse even when full-prefill quality remains stable, and that reusable-context length affects the latency benefit. We also study reuse of independently prepared skill caches on SkillsBench tasks. HYPIC TRR retains 83.0% of the reward improvement provided by skills, while reducing first-call time-to-first-token to roughly 20–30% of full prefill on executions with large reusable-context fractions. These findings establish cache preparation and request structure as important evaluation dimensions and provide a common framework for studying the benefits and limitations of approximate KV-cache reuse.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.