Cache Engineering for Agents
Abstract
Large Language Model (LLM) agents are rapidly evolving to handle complex tasks, navigating long sequences of observations and actions. However, as agent trajectories grow longer, Key-Value (KV) cache storage and I/O become dominant costs. Long KV caches can also degrade model performance, especially when an LLM's context window limit is exceeded. Prior work uses context engineering to reduce cache costs and support agent trajectories longer than the LLM’s context window. They typically trim or summarize information irrelevant to the current round, thereby reducing the cache used in subsequent attention computations. However, modifying the prefix invalidates all subsequent cached states and triggers costly re-computation of the cache for the remaining context. Compact summarizations can also directly lose critical information, reducing accuracy in long-horizon tasks. To mitigate these limitations, we propose a new cache engineering scheme. It lets agents proactively drop selected KV cache while directly reusing retained ones, without relying on context engineering as a proxy. The key challenge is the history dependence of cache states: the same token sequence can yield different caches under different cache-operation histories, making token-prefix matching insufficient to determine whether a cache can be reused. Thus, we propose AKV, an efficient serving engine to support cache engineering. It provides a simple, stateless interface compatible with existing API-based serving, and hides the complexity of stateful cache management inside. We further extend the radix tree and attention masking to maximize cache reuse, enable timely eviction, and efficiently restore missing cache states. Experiments show that AKV enables a simple rolling-drop cache engineering strategy that effectively improves benchmark scores by 1.7 points while reducing the average cost by 27.7% across various LLMs and agent benchmarks. It further enables recursive self-improvement of agents' cache-engineering strategy, expanding the cost reduction to 17.8%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.