acceptodds
Under review as a conference paper at ICLR 2027

Supervision Interfaces for Privileged Coordination Advice in Cooperative MARL

Abstract

Cooperative multi-agent reinforcement learning commonly adopts the Centralized Training with Decentralized Execution (CTDE) paradigm, in which global infor- mation is used during training while execution relies only on local observations. However, global information typically influences the actor indirectly through a centralized critic, leaving unclear how privileged coordination advice should be delivered to decentralized actors. We propose that privileged advice can supervise behavior directly through imitation or policy regularization, or supervise repre- sentation while reinforcement learning remains responsible for action selection, allowing the actor to absorb privileged coordination structure without directly im- itating the advised behavior. We introduce Factorized Privileged Action Super- vision (FPAS), which maps privileged coordination decisions to per-agent targets and trains each decentralized actor to predict them through a train-only auxiliary head. FPAS requires no learned teacher and adds no execution-time overhead. On SMACv2, FPAS outperforms MAPPO across all three races (+8.2/ + 12.8/ + 5.3 pp; paired p < 0.03 in each race, with 15/15 paired seed wins), with two races held out from tuning. Matched control experiments show that the gains are associated with the prediction pathway and decision-relevant target content. A loss-matched intervention further sharpens this conclusion: the same advice and cross-entropy degrade learning when applied through policy logits, but recover the FPAS gain when applied through the auxiliary prediction head. Target interventions further show that consistency, local inferability, and decision content are important prop- erties of effective representation supervision. Across MPE, SMACv2, and Over- cooked, imitation is strongest when the advice defines a strong and locally attain- able policy, whereas prediction is more reliable when the advice provides only partial coordination structure or when direct execution is constrained. These re- sults show that the supervision route through which privileged information enters the actor is itself an important design variable in CTDE.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.