acceptodds
Under review as a conference paper at ICLR 2027

A Framework for Enhancing Multi-agent Cooperation through Communication Policy Transfer

Abstract

By enabling agents to exchange complementary information, communication-based multi-agent reinforcement learning (MARL) substantially enhances coordination, spurring growing interest in transferring its cooperative knowledge to decentralized agents under centralized training with decentralized execution (CTDE). However, communication policies may exploit information unavailable during decentralized execution, causing policy discrepancy and negative transfer. We theoretically analyze this gap and derive a Q-smoothing condition requiring the communication Q-function to remain consistent across feasible observations of other agents under the same local observation. Based on this analysis, we propose Communication Smoothing Distillation (CSD). CSD learns cross-agent conditional observation distributions and constructs latent communication through distribution sampling. Gradient ascent is further employed to search for communication variations that maximize Q-value discrepancy, while a smoothing perturbation loss constrains Q-value. The communication-smoothed agents provide Q-value targets for CTDE students, enabling the transfer of cooperative knowledge while reducing policy discrepancy. Extensive evaluations on partially observable cooperative tasks demonstrate that CSD mitigates negative transfer, outperforms advanced communication policy transfer and CTDE methods, and can be integrated into existing value-decomposition algorithms.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.