acceptodds
Under review as a conference paper at ICLR 2027

Multi-Agent Reinforcement Learning with On-Policy Distillation

Abstract

Centralized training with decentralized execution (CTDE) exploits privileged information during training, yet decentralized multi-agent reinforcement learning policies often remain driven by sparse environmental rewards. Online policy distillation can provide dense supervision, but applying it to MARL introduces observation mismatch between a globally informed teacher and local students, as well as policy learning mismatch between their continuously evolving policies. We propose MAOPD, an online policy distillation framework that jointly trains an autoregressive centralized teacher and decentralized parallel students. Agent distillation bridges observation mismatch by using a local prefix predictor to construct locally realizable decision representations that match privileged teacher features. Action distillation addresses the structural policy mismatch by factorizing the teacher's joint policy into agent-wise conditional targets on student-generated trajectories. The teacher and students further share most policy parameters and co-adapt online through separate reinforcement-learning objectives, alleviating evolving policy mismatch without separate teacher pre-training. Across 26 tasks from multiple MARL benchmarks (i.e., RWARE, SMAX, and MPE), MAOPD achieves the highest aggregate performance under matched environment-interaction budgets and ranks first or tied for first on 22 tasks. Ablations confirm the complementary benefits of both distillation objectives and the online co-adaptation mechanisms.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.