acceptodds
Under review as a conference paper at ICLR 2027

Harness-Agnostic Predictive KV-Cache Prefetch Engine for LLM Agents

Abstract

Long-running LLM agents alternate between model inference and tool execution, creating idle periods during which cached prefixes may be evicted. Subsequent requests can also encounter uncached tool outputs or cold-started subagents. Because inference servers do not observe tool execution, returned results, or subagent creation before the next request arrives, they cannot use these gaps to prepare the KV cache. We introduce a harness-agnostic predictive KV-cache prefetch engine that uses agent-runtime lifecycle events to move KV-cache preparation off the critical path. During tool execution, the engine issues low-priority warmups for the known prefix; after completion, it warms the anticipated next-request prefix with the actual tool result. It also prepares initial prefixes for newly spawned subagents. A contextual Thompson sampling policy selects whether and how to prefetch based on cache state, prefix length, and predicted tool duration, balancing cache reuse against warmup cost and foreground interference. Across Claude Code, Hermes, and Codex on GAIA and SWE-bench Lite, the engine reduces aggregate time to first token by 67.6%, 19.6%, and 68.9%, respectively. Subagent warming reduces median first-call TTFT by 68.6%, from 564 ms to 177 ms. At eight concurrent sessions, warmups add 0.5% prompt-token overhead. These results demonstrate that harness-visible execution gaps can be used to prepare anticipated KV state and reduce cold-prefill latency in long-running agents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.