acceptodds
Under review as a conference paper at ICLR 2027

Coordination Scaffolds: Task-Dependent Withdrawal of Privileged Observations in Cooperative Multi-Agent Reinforcement Learning

Abstract

Training-only privileged observations can simplify reinforcement learning, but a deployable policy must eventually act without them. Existing guided-observability methods typically use one withdrawal recipe across tasks. We show that this assumption breaks in cooperative multi-agent reinforcement learning (MARL), where privileged state can scaffold the discovery of coordination itself. Across SMAC and MPE, the same withdrawal schedule can improve one task and degrade another: augmentation prevents bootstrap failure on the hardest map, while linear decay loses up to 0.205 win rate on another and fast decay is consistently poor on low-demand MPE tasks. Schedule-mismatch interventions, dose-matched timing controls, and withdrawal-recovery analyses show that these effects arise along the optimization path; the final removal produces little immediate disruption. We introduce Observation-Control Irreducibility (), a probe-conditioned diagnostic based on the predictive gain from modeling agents' local observations jointly rather than additively. An OCI-guided schedule uses this demand estimate to choose earlier or later withdrawal. It avoids the large fixed-schedule failures on the development tasks and, when frozen before training, is oracle-best on one informative held-out task and within 0.078 of the oracle on another. A QMIX replication preserves the task-dependent diagnosis but requires a different demand-to-schedule calibration. These results identify privileged observations as coordination scaffolds whose withdrawal should be conditioned on the task and learner.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.