Latent Communication Between LLM Agents: Channels, Alignment, and the Limits of Text
Abstract
Multi-agent LLM systems typically communicate via text, but language models may encode semantic structure that resists faithful serialisation into tokens. We construct three communication channels between LLM agents — a lossless dense latent channel, a sparse channel built on pretrained Sparse Autoencoder (SAE) features, and standard text — and measure how much concept-discriminating information each preserves. The SAE-sparse channel retains 99.4% probe accuracy at 28× compression over the dense channel, versus 80.4% for text. Text serialisation does not attenuate SAE features but replaces them: 88% of a concept’s active features are lost after a text round-trip, and we show that these lost features encode surface form rather than task-relevant semantics — re-injecting them degrades performance. Cross-architecture Procrustes alignment between Llama 3.1 8B and Mistral 7B achieves 92% top-1 concept retrieval, yet on every task type tested the text channel matches or exceeds the latent channel by 3–10 percentage points. We report an honest negative result: the infrastructure is sound and architectures do converge, but no task in our current suite requires information text cannot express — pointing future work toward richer tasks rather than better channels.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.