acceptodds
Under review as a conference paper at ICLR 2027

ForeCache: Efficient LLM-Based Multi-Agent Systems with Intent-Conditioned Pruning and Amortized KV Reuse

Abstract

Multi-agent systems based on large language models (LLMs) address complex tasks such as code generation by exchanging and refining intermediate outputs, but repeatedly processing accumulated interaction histories leads to high inference latency. Existing approaches reduce this overhead through history compression or key-value (KV) cache reuse. However, compression can remove information needed by the next agent, while direct KV reuse can reduce downstream answer accuracy when cached states no longer match that agent's context. To address these limitations, we propose ForeCache, a training-free framework for efficient multi-agent inference. Specifically, consumer-conditioned saliency (CCS) uses each downstream agent's instructions to estimate the task relevance of shared history and guide pruning. We also introduce divergence-guided recomputation (DGR), which reuses upstream KV states and stops further recomputation for tokens whose cached and updated states are sufficiently close, reducing repeated history encoding. Experiments on well-known benchmarks show that ForeCache maintains task performance comparable to full prefilling. In the six-agent setting, it achieves up to a 6.6× speedup in the final agent's time to first token (TTFT) over full prefilling.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.