MMA: Efficient Latent Communication for Parallel Large Language Model Agents
Abstract
Large Language Models (LLMs) process context through a sequential token interface. However, current multi-agent systems naturally involve parallel branches with disjoint evidence or private contexts. Existing agent communication methods through text or sequence-valued latent space incur decoding re-prefill overhead and introduce communication costs that grow with context length. In this paper, we propose Multi-Modal Attention (MMA), a latent communication framework for frozen parallel LLMs. Each branch retains its context and Key Value (KV) cache locally, exchanging only current-token hidden states at sparse synchronization layers. Furthermore, a latent-fusion aggregator produces a shared next-token distribution for all branches. Because MMA communicates current-token states rather than sequence-length-dependent memories, its payload is independent of input length, and the same interface naturally extends to multiple branches, heterogeneous model widths, and different input modalities. Evaluated on information retrieval benchmarks spanning different context lengths, MMA matches or outperforms other communication baselines while remaining efficient with significantly less communication payload. Its performance advantages are most pronounced in long-context, multi-branch retrieval. These results support MMA as a scalable communication interface for parallel multi-agent inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.