A Framework for Enhancing Multi-agent Cooperation through Communication Policy Transfer
Abstract
By enabling agents to exchange complementary information, communication-based multi-agent reinforcement learning (MARL) substantially enhances coordination, spurring growing interest in transferring its cooperative knowledge to decentralized agents under centralized training with decentralized execution (CTDE). However, communication policies may exploit information unavailable during decentralized execution, causing policy discrepancy and negative transfer. We theoretically analyze this gap and derive a Q-smoothing condition requiring the communication Q-function to remain consistent across feasible observations of other agents under the same local observation. Based on this analysis, we propose Communication Smoothing Distillation (CSD). CSD learns cross-agent conditional observation distributions and constructs latent communication through distribution sampling. Gradient ascent is further employed to search for communication variations that maximize Q-value discrepancy, while a smoothing perturbation loss constrains Q-value. The communication-smoothed agents provide Q-value targets for CTDE students, enabling the transfer of cooperative knowledge while reducing policy discrepancy. Extensive evaluations on partially observable cooperative tasks demonstrate that CSD mitigates negative transfer, outperforms advanced communication policy transfer and CTDE methods, and can be integrated into existing value-decomposition algorithms.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.