SDM: Synergistic Dependency Maximization for Cooperative Multi-Agent Reinforcement Learning
Abstract
Cooperative multi-agent reinforcement learning requires agents to develop complementary behaviors whose task value emerges through interaction with their teammates. A fundamental challenge is to quantify task-relevant synergy arising from these interactions and translate it into an effective signal for policy learning. To address this challenge, we introduce Synergistic Dependency Maximization (SDM), which promotes task-relevant interaction coupling by encouraging an agent's action to become more informative about the multi-step team return when its teammates' actions are taken into account. Using Partial Information Decomposition (PID), we formalize synergistic dependency through synergy beyond redundancy, yielding a principled task-relevant coordination criterion for multi-agent interactions. To estimate this dependency reliably, SDM adopts a conservative estimation strategy and uses the resulting signal to guide policy optimization while retaining the original task objective for value estimation. Experiments across MPE, MAMuJoCo, SMAC, and SMACv2 demonstrate consistent improvements over strong multi-agent policy optimization baselines and existing explicit coordination objectives. These results establish synergistic dependency as a principled objective for promoting cooperative behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.