CacheCodec: Task-Oriented KV Compression for Bandwidth-Constrained Multi-LLM Communication
Abstract
Large language models are increasingly deployed in collaborative systems, where exchanging key-value (KV) states can bypass autoregressive intermediate-text generation and preserve rich model-internal context. However, KV caches are high-dimensional, model-specific states designed for local inference rather than communication, making direct transmission prohibitively expensive over bandwidth-constrained links. We therefore formulate KV compression as task-oriented communication, preserving downstream-relevant information rather than faithfully reconstructing the transmitted KV states. Based on this perspective, we propose CacheCodec, a task-oriented codec for efficient inter-LLM KV communication. CacheCodec first maps high-dimensional KV states into a compact representation message optimized under downstream task supervision, and then applies frequency-domain coding along the sequence dimension to convert this representation into a low-rate bitstream for transmission. Across heterogeneous LLM pairs and downstream tasks, CacheCodec reduces transmitted payload by while maintaining or improving task accuracy, and achieves reduction in end-to-end latency over a 50 Mbps communication link.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.