Receiver-Conditioned Selection and Reconstruction for Latent Communication Between LLM Agents
Abstract
Latent communication through key-value (KV) caches can avoid natural-language decoding and repeated prefilling by transferring internal states between large language model (LLM) agents. Existing approaches fix transmission choices before the current query is known or reuse latent states formed outside the receiver context. We find that the content a sender ranks first can hurt the receiver and that different receivers need different content. We therefore condition what to retrieve and how to read it on the receiver. Each document is written once, before any receiver or query exists, as an INT8 snapshot of its early-layer state. At read time, the receiver uses its state to select a document and reruns the upper layers under its private context, so document states are formed for that receiver rather than transplanted. Across Qwen3-4B, Qwen3-8B, and Llama-3.1-8B-Instruct, receiver-aware selection with BGE-M3 improves F1 by 15.5–17.2 points over query-only selection for both text and checkpoint delivery. With selection fixed, read-time reconstruction improves F1 by 5.1–13.5 points over full static KV reuse with payloads to smaller. Under controlled TCP networking, the complete system reduces time to first token (TTFT) by 6.91%, 5.81%, and 4.23% on Qwen3-4B, Qwen3-8B, and Llama-3.1-8B-Instruct, respectively. In the end-to-end evaluation, answer quality remains comparable to native text, with F1 differences ranging from to points. Code is available at anonymous.4open.science/r/rcsr-latent-comm-B184.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.