Racer: Recoverable KV Cache Management for Multi-Turn LLM Agents
Abstract
Efficient context management is critical for multi-turn large language model (LLM) agents, whose tool definitions and growing interaction histories increase key-value (KV) cache usage. Recent work has explored text-level context compression and KV cache compression to reduce this cost. However, compression may lead to a decision that differs from full-context execution, which in multi-turn agents can change the environment and affect subsequent decisions through the resulting observation. We refer to this phenomenon as compression-induced trajectory drift. To address this problem, we propose Racer, a recoverable KV cache management framework compatible with different KV compression methods. For tool definitions, Racer retains query-relevant tool definitions with full KV representations while compressing the remaining definitions. For interaction histories, Racer protects selected historical information as native KV before generation, then uses a risk detector to identify potentially incorrect drafts and selectively restores historical evidence, allowing the model to regenerate its decision before execution. Extensive evaluations demonstrate that Racer improves mean task success by 43.7% across six KV-compression backends on long-context tasks, while achieving 1.67x higher task throughput than full-context execution under concurrent serving.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.