On the Role of Communication in Offline Multi-Agent Reinforcement Learning
Abstract
This paper presents CODA, a framework for learning inter-agent communication from offline multi-agent datasets collected without communication. CODA jointly learns action and communication policies from fixed offline data, without message supervision or additional environment interaction, and deploys the learned communication after training. We analyze the role of communication in offline multi-agent RL, where communication can provide additional information about dataset-supported actions but can also increase the difficulty of learning a message-conditioned actor. Our analysis formalizes this trade-off and provides a condition under which communication tightens an overestimation bound. Experiments across several multi-agent tasks show that CODA can substantially improve offline policy learning, with gains that depend on the underlying offline MARL algorithm and dataset. Further analysis shows that communication often reduces policy-data action deviation and, in several settings, empirical critic overestimation, consistent with the stabilization perspective suggested by the theory. CODA also remains effective under delayed communication in several tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.