HorizonKV: Future-Aware KVCache Reuse via Lightweight Relevance Estimation for Interleaved Agent Contexts
Abstract
In agent scenarios, context interleaving continuously combines historical context with new context, including user questions, tool calls, and more. KVCache reuse reduces prefill computation, while recomputing important positions maintains accuracy. However, existing methods select recomputation positions based on fixed positions or historical relevance, which fail in dynamically evolving agent contexts. We observe that the next token query helps identify historical positions relevant to subsequent generation. This motivates future-aware KVCache reuse for interleaved contexts, which introduces two challenges. First, this query is unavailable when recomputation positions are selected, and obtaining it requires full prefill. Second, even after selecting recomputation positions, their recovery only requires a subset of historical dependencies, introducing redundant computation. To address these challenges, we propose HorizonKV, a two-stage sparse recovery framework guided by future relevance estimated from current dynamic content. First, we propose future relevance guided recomputation selection, which averages shallow-layer queries at dynamic positions to score historical positions and select them for recomputation within the budget. Second, we propose future relevance guided shared-key recovery, which reuses these scores to construct a shared KV set, reducing recovery computation without additional dense query-key scoring. On agent traces, HorizonKV improves the average score by up to 64.9% over CacheBlend under the same 10% recomputation ratio. On HotpotQA agent trace, it achieves a 5.85× TTFT speedup over CacheBlend at comparable accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.