AttuneKV: Receiver-Aware KV-Cache Handoff for Efficient Multi-Agent Inference
Abstract
Multi-agent large language model (LLM) systems frequently process messages generated by other agents. When agents share a model, the sender’s key–value (KV) cache can be transferred to avoid redundant prefilling, but the imported cache is context-dependent and may not match the receiver’s context. Moreover, the receiver may need only part of the message during decoding. KV handoff therefore poses two receiver-side resource decisions: where to spend recomputation and which transferred entries to keep in memory. We present AttuneKV, a receiver-aware KV-cache handoff method that obtains both decisions from one lightweight receiver observation. AttuneKV forwards one existing receiver-side token over the imported cache. Its layerwise attention ranks message positions for recomputation, while the attention-output change induced by removing each entry ranks positions for retention. The same observation supports both decisions without forwarding the incoming message positions for selection or maintaining a persistent KV reference pool. At 12,288 incoming tokens, AttuneKV achieves a 2.92× receiver time to first token (TTFT) speedup over dense prefill and a 1.45× speedup over RelayCaching. It also reduces KV-handoff memory by 2.0× and 34× relative to dense prefill and KVCOMM, respectively. Across math problem solving, broad-domain knowledge evaluation, and code generation, AttuneKV remains within 3.92 pp of dense prefill performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.