acceptodds
Under review as a conference paper at ICLR 2027

BEAD: Blind-spot Entity-aware Advantage Distillation for Cooperative Multi-Agent Reinforcement Learning

Abstract

Centralized Training with Decentralized Execution (CTDE) is widely adopted in cooperative Multi-agent Reinforcement Learning. Building on CTDE, Centralized Teacher with Decentralized Student (CTDS) exploits global information to construct teachers, but indiscriminate use of such information may degrade teacher decision quality. We propose Blind-spot Entity-aware Advantage Distillation (BEAD), a centralized-teacher decentralized-student framework that selectively exploits global information for effective knowledge transfer. BEAD represents the global state as entities and employs gated blind-spot compensation to selectively incorporate locally unobserved entities into the privileged teacher. In addition, BEAD introduces explicit advantage distillation anchored to the student's value estimates, transferring the teacher's relative action preferences while preserving the student's own value learning. Theoretical analysis shows that BEAD preserves action-selection preferences, while experiments on challenging multi-agent reinforcement learning benchmarks demonstrate consistent improvements over strong CTDE and teacher-student baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.