AgentKV: Predictive Admission Prevents KV-Cache Collapse in Agentic LLM Serving
Abstract
Long-horizon language-model agents place unique demands on LLM serving systems because they repeatedly return from tool calls with increasingly long contexts. At high concurrency, these growing contexts can evict one another’s reusable KV cache, forcing repeated prefill, slowing completion, and keeping even more contexts resident. We call this positive-feedback failure KV-cache collapse. We show that the onset of collapse can be predicted from the relationship between available KV capacity and the average context demand of each instance. We introduce AGENTKV, a predictive admission controller that forecasts shorthorizon context growth, maintains a lifecycle-safe working set under an adaptive high-bandwidth memory (HBM) budget, and replans from live cache telemetry. Across ten Qwen3-8B and Qwen3.8-27B campaigns on four RTX PRO 6000 GPUs, the predicted capacity boundary closely matches the observed transition from stable prefix reuse to cache collapse. Beyond this boundary, AGENTKV outperforms vLLM, ThunderAgent, InferCept, Continuum-TTL, and cache-aware vLLM-router in every measured cell, reaching up to 3.7× the throughput of stock vLLM. At 240 concurrent instances, AGENTKV retains roughly 75% prefix reuse for Qwen3-8B workloads while stock vLLM falls to at most 0.4%, and removing predictive admission leaves only 27–77% of full-system throughput. These results show that predictive working-set control can prevent a high-load cache-collapse regime that allocation, routing, and persistence alone do not fully address.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.