Can a Main Agent Reuse Its Sub-Agents' KV Caches? A Systematic Empirical Study
Abstract
Multi-agent systems incur substantial inference costs, yet agent harnesses typically operate without awareness of the underlying inference engine. We investigate a simple question: Can a main agent reuse the KV cache of its sub-agents? We build a cross-layer system that enables this reuse and supports an empirical study spanning content reconstruction, retrieval-augmented question answering, and multi-agent trajectories. Our study suggests that conclusion-token KV states retain information about both the conclusion and its preceding context. Reusing these states can preserve answer quality in document question answering and limited-fan-out agent workflows while avoiding redundant prefill computation. However, accuracy degrades as the number of contributing sub-agents increases, exposing limits to direct cache reuse. These findings identify both opportunities and challenges for cross-agent KV sharing, motivating a broader shift toward co-designing agent harnesses and LLM inference systems, with workload-aware coordination across their traditional boundary.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.