Can LLM Agents Communicate Without Autoregression? Thought Blocks for Latent Multi-Agent Collaboration
Abstract
LLM-based multi-agent systems typically communicate through text, incurring repeated generation and re-encoding costs. Latent communication removes explicit verbalization, but constructing a latent message may still require a sequential latent trajectory. We therefore distinguish message representation from message construction and ask: can an agent directly construct a useful continuation state without executing the trajectory that normally produces it? We formulate this as trajectory-to-state compilation and instantiate it with Thought Blocks, fixed-size blocks of latent key–value (KV) memory constructed in one shot. Thought Blocks map a compact continuation representation and compressed contextual KV memory to parallel latent positions, which are rendered into model-native KV continuation states in a single frozen-backbone forward pass. We train the writer by distilling continuation states reached by an autoregressive latent teacher, whose trajectory is omitted at inference. Across multiple LLM backbones and reasoning and question-answering tasks, Thought Blocks substantially reduce sequential communication computation and latency relative to autoregressive latent communication while maintaining competitive accuracy, achieving a 2.5 times average latency speedup. Compared with text-based multi-agent communication, the efficiency gains are broader still, reducing text-token traffic, autoregressive computation, and latency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.