Recoverable Coordination Distillation for Decentralized Multi-Agent Reinforcement Learning
Abstract
Centralized training with decentralized execution (CTDE) allows multi-agent learners to exploit privileged information during training, but deployed agents must act from local inputs. Existing approaches may attempt to reconstruct missing teammate information, although much of this information can be irrelevant to the current decision. We propose Recoverable Coordination Distillation (RCD), which instead predicts the locally recoverable effect of privileged information on a teacher's action preferences. A frozen teacher defines this coordination effect as the centered difference between its action logits with and without privileged input. A local predictor estimates the conditional mean and uncertainty of this effect, while a reliability-aware fusion mechanism regulates its influence on the decentralized policy. At deployment, all computations use only execution-available local information, without the teacher, privileged inputs, or communication. We characterize the locally recoverable component of the teacher effect and derive a return bound relating prediction errors to policy performance. Across seven tasks from MPE, LBF, and RWARE, RCD achieves higher mean final return than our MAPPO-based MAE baseline in all task–input comparisons. With local history, it also achieves higher mean learning-curve AUC than the Local baseline on all seven tasks. These results support recovering decision-relevant coordination effects rather than reconstructing missing information for decentralized multi-agent control.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.