KVTrans: Towards Seamless Cross-Context KV Cache Reuse for Efficient Agentic Communication
Abstract
LLM-based multi-agent systems repeatedly re-encode received messages, even though their key–value (KV) states have already been computed during sender generation. Reusing these states can reduce redundant computation, but differences in agent instructions, interaction histories, and token positions make direct cross-context reuse unreliable. We introduce **KVTRANS**, a **training-free** framework for **adapting sender-generated KV caches to receiver contexts** among agents sharing the same model. KVTRANS combines rotary position alignment, receiver-referenced key and value rescaling, and dynamic modulation of receiver-instruction keys during decoding. Across four models on MMLU debate, KVTRANS improves macro-averaged accuracy from 59.61% under direct reuse to 65.44%, compared with 64.88% for text communication, while reducing receiver-side prefill tokens by **59.14%**. Timing measurements on four matched role pairs show a **57.04%** reduction in prefill time, including pre-decoding cache-adaptation overhead. In WebSailor experiments on BrowseComp-EN and BrowseComp-ZH, KVTRANS partially recovers the quality degradation caused by direct reuse while processing only **53.9%** and **62.3%** of the text baseline’s receiver-side prefill tokens, respectively. These results demonstrate that lightweight receiver-side adaptation can improve the quality–efficiency trade-off of cross-context KV reuse without retraining model weights.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.