KVISE: selective KV sharing for efficient distributed LLM collaboration
Abstract
Communication overhead can become a major bottleneck in collaborative LLM inference when agents exchange large intermediate states. Although KV-cache sharing enables direct reuse of sender-side computation, exhaustively transferring remote caches is often inefficient because KV utility is highly non-uniform and depends on the receiver's current computation. We present KVise, a training-free and plug-and-play framework that formulates inter-agent KV exchange as receiver-conditioned communication allocation. The receiver provides compact native attention queries, while senders summarize candidate KV shards using mergeable online-softmax statistics, enabling their utility to be estimated before transmitting raw tensors. KVise then selects shards according to their marginal attention-distortion reduction per byte and jointly reuses the selected states with the receiver's local cache. Across four model families and four multi-hop QA datasets, KVise reduces logical wire traffic by 30.2–61.2% and mean TTFT and completion latency by approximately 21% under our main network setting. Further analyses show concentrated receiver-conditioned KV utility, a controllable quality–communication trade-off, and diminishing latency gains as communication becomes less dominant. These results demonstrate that selective, receiver-aware KV transfer can substantially reduce unnecessary inter-agent communication while preserving flexible control over efficiency and task quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.