Rent or Recompute?: Learning-Augmented Placement of Agent KV Caches Across Memory and Shared Storage Tiers
Abstract
LLM agents keep contexts of 100K tokens or more alive across dozens of turns. Most of that time, they sit idle, waiting on tool calls, sub-agents, and people. Each idle gap forces a serving system to decide whether a session's KV cache is worth keeping and where it should live. Production traces show that prefix-cache hit rates collapse once a session idles for minutes, and a miss costs several times the pre-fill for the next turn. Today, this decision is made based on recency or fixed time-to-live rules. We show that placing an idle session's KV cache across GPU memory, host memory, and storage is an instance of multi-state dynamic power management. Each tier is a state with a rent rate. Returning from a tier, by reload or by recomputation, is its wake-up cost. This lets us import learning-augmented ski-rental guarantees directly: a policy that consumes a predicted idle gap is -consistent and -robust per gap, for a trust parameter . We train an idle-gap predictor on features observable when a request finishes, and evaluate it on 346K idle gaps from 4,265 real coding-agent sessions, split by user. Whether a gap ends in a tool result or a human reply is observable, and it matters (seconds versus minutes). The length of the gap within each kind is only weakly predictable: 42% of tool gaps and 24% of human gaps fall within of the prediction, against 26% and 23% for a per-tool lookup table. On held-out users, the learning-augmented policy meets its per-gap bounds exactly. At a reference parameter set, a hedged policy with a validation-tuned trust parameter uses 11% fewer GPU-seconds than the best prediction-free policy; exact predictions would save 20%. Using the same predictions without hedging costs the clairvoyant optimum. A closed-form crossover context length, , says when a storage tier beats recomputation. We measured its inputs with vLLM for an 8B model on H100, A100, and L40 GPUs, and for a 72B model on four H100S GPUs, across five kinds of storage paths from seven clients on three clusters. The binding quantity is usually the client's network link, not the storage system: tiers cap at the link (1.2 GB/s on 10 GbE nodes, 3.0 GB/s on 25 GbE). At measured prefill costs, an L40 breaks even against every mount we measured; an A100 against a 10 GbE NFS mount at 34K tokens; and an H100 against a 25 GbE mount at 32K tokens, or a node-local NVMe at every length. The 72B model prefills slower than the 8B model because tensor parallelism over PCIe returns only a third of the four GPUs. It therefore needs only 1.1 GB/s, and a 1.3 GB/s NFS mount beats recomputing it at every length. Re-running the policy comparison with the measured parameters, where prefill is 4–28 costlier per token than the assumed reference set, removes the learned policy's advantage (0-3%). With a 1 s host-memory round trip against a 9-32s re-prefill, the decision points move into the tails, where the predictor has no skill. Our analysis identifies this as the regime where prediction has the least to decide, and the advantage returns (10%) only when host memory becomes expensive. Because the policy consumes only the predicted tier, the measurements turn into targets: a predictor must flag the long tool calls that a regression on gap length misses, and a link must deliver 2.5 GB/s to an H100 for storage to pay. We release the harness, simulator, and predictor as a testbed for both.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.