acceptodds
Under review as a conference paper at ICLR 2027

LLM-SemanticLAN: Anchored Public Latent Communication via KV Caches across Heterogeneous Large Language Models

Abstract

Latent communication is the key technique to coordinate heterogeneous large language models for composite tasks, which directly exchanges information between internal states such as KV caches instead of texts to avoid slow sequential decoding and redundant prefilling. However, existing latent communication either maps caches between one pair of models at a time or relies on a shared space defined only by adapters trained together, leaving no fixed coordinates that a new model of another scale, family, architecture, or modality could align to. To address these issues, we formalize Public Latent Communication (PLC), which links every model via a shared latent space like a local area network (LAN). We realize PLC in LLM-SemanticLAN, a lightweight and high-efficiency framework that anchors the coordinates of this space in closed form by ridge regression before training, using targets from a frozen reference model. LLM-SemanticLAN employs model-specific codecs to map native KV caches of LLMs into the compact shared representations of this space for transfer and reconstruction, with codecs of only 1.0%-3.6% of the base model parameters and latent representations 64× smaller on average than the native KV caches at the same BF16 precision. Extensive evaluations across scales, attention architectures, families, and modalities verify its strong generalization, showing that aligning a new model with only a single existing member in the shared latent space, supervised only by the receiver's behaviour, enables zero-shot communication across all unpaired members. Without benchmark-specific training, LLM-SemanticLAN achieves near-lossless parity with text-to-text communication and superior vision-to-text performance, while delivering 12-35× speedups in end-to-end latency and 2.2-4.3× FLOPs reductions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.