BEAD: Blind-spot Entity-aware Advantage Distillation for Cooperative Multi-Agent Reinforcement Learning
Abstract
Centralized Training with Decentralized Execution (CTDE) is widely adopted in cooperative Multi-agent Reinforcement Learning. Building on CTDE, Centralized Teacher with Decentralized Student (CTDS) exploits global information to construct teachers, but indiscriminate use of such information may degrade teacher decision quality. We propose Blind-spot Entity-aware Advantage Distillation (BEAD), a centralized-teacher decentralized-student framework that selectively exploits global information for effective knowledge transfer. BEAD represents the global state as entities and employs gated blind-spot compensation to selectively incorporate locally unobserved entities into the privileged teacher. In addition, BEAD introduces explicit advantage distillation anchored to the student's value estimates, transferring the teacher's relative action preferences while preserving the student's own value learning. Theoretical analysis shows that BEAD preserves action-selection preferences, while experiments on challenging multi-agent reinforcement learning benchmarks demonstrate consistent improvements over strong CTDE and teacher-student baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.