ForeCache: Efficient LLM-Based Multi-Agent Systems with Intent-Conditioned Pruning and Amortized KV Reuse
Abstract
Multi-agent systems based on large language models (LLMs) address complex tasks such as code generation by exchanging and refining intermediate outputs, but repeatedly processing accumulated interaction histories leads to high inference latency. Existing approaches reduce this overhead through history compression or key-value (KV) cache reuse. However, compression can remove information needed by the next agent, while direct KV reuse can reduce downstream answer accuracy when cached states no longer match that agent's context. To address these limitations, we propose ForeCache, a training-free framework for efficient multi-agent inference. Specifically, consumer-conditioned saliency (CCS) uses each downstream agent's instructions to estimate the task relevance of shared history and guide pruning. We also introduce divergence-guided recomputation (DGR), which reuses upstream KV states and stops further recomputation for tokens whose cached and updated states are sufficiently close, reducing repeated history encoding. Experiments on well-known benchmarks show that ForeCache maintains task performance comparable to full prefilling. In the six-agent setting, it achieves up to a 6.6× speedup in the final agent's time to first token (TTFT) over full prefilling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.