Emergence, Not Bandwidth: Physical Coupling and the Limits of Learned Multi-Agent Communication
Abstract
Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message should encode, what compression costs over a horizon, and when a learned protocol is unique enough to be read by a teammate. We answer them for rate-limited Dec-POMDPs, and then measure how far reinforcement learning falls short of the optimum the theory locates. Our theorems fix what is achievable independently of any learner, so a gap between an engineered sender and a learned one at the same bit budget is an optimization fact rather than an information-theoretic one. We instantiate this on three MuJoCo continuous-control arenas spanning zero, partial and rigid physical coupling, with every condition charged exactly 2 bits per decision by construction, and the discriminating regime is produced by closing a physical side channel within one arena, holding bodies, task and reward fixed. Communication value is governed by coupling: where the agents are rigidly coupled through a shared object, no channel beats silence (, , ), because proprioception already carries what a message would say; where they are uncoupled, every condition solves the task; and in the partially coupled regime the engineered -bit sender reaches an interquartile mean of while the learned -bit sender reaches and is indistinguishable from silence (, ). Because the two senders hold the same alphabet, bandwidth cannot explain the gap. Initializing a learned run from an engineered run's receiver localizes the failure: the same channel then reaches against cold-started (), so the failure is neither representational nor a matter of maintenance: what reinforcement learning cannot do here is discover the protocol in the first place. Cross-play shows the learned protocols are individually meaningful and mutually unintelligible: self-play collapses to across seeds, and the best alignment we can construct leaves at least of that gap standing. Every headline result is reported at seeds per arena, against seven published baselines reimplemented in our trunk at matched rate.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.