Multi-Agent Information-Guided Policy Optimization: An Information-Theoretic Perspective
Abstract
In multi-agent reinforcement learning (MARL), partial observability limits the ability of decentralized agents to leverage global information, motivating widely used frameworks such as Centralized Training with Decentralized Execution (CTDE) and more recently, Multi-Agent Guided Policy Optimization (MAGPO). However, MAGPO relies solely on policy alignment and fails to explicitly align decision-relevant information between the guider and learner, limiting decentralized policy learning. In this work, we revisit the imitation gap between guider and learner from an information-theoretic perspective, and show that it fundamentally arises from an information gap—a mismatch in decision-relevant information between centralized and decentralized policies. To mitigate this gap, we propose Multi-Agent Information-Guided Policy Optimization (MAIGPO), which employs the information bottleneck principle to extract decision-relevant representations and mutual information maximization to align them between the guider and learners through bidirectional information flow. By encouraging learners to capture the guider’s decision rationale and constraining the guider to learner-accessible information, MAIGPO significantly improves decentralized policy learning and outperforms state-of-the-art methods on standard benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.