Descent-Guided Policy Gradient for Scalable Cooperative Multi-Agent Learning
Abstract
Scaling cooperative multi-agent reinforcement learning (MARL) is fundamentally limited by cross-agent noise. When agents share a common reward, each agent's learning signal is computed from a shared return that depends on all agents, so the stochasticity of the other agents enters the signal as cross-agent noise that grows with . Many engineering systems, such as cloud computing and power systems, have differentiable analytical models that prescribe efficient system states, providing a new reference beyond noisy shared returns. In this work, we propose Descent-Guided Policy Gradient (DG-PG), a framework that augments policy-gradient updates with a low-noise descent signal toward this reference. We prove that DG-PG reduces policy-gradient estimator variance from to , preserves the stationary points of the cooperative objective, and achieves agent-independent sample complexity . We evaluate DG-PG on cloud resource scheduling and power dispatch tasks, which involve discrete and continuous decisions, respectively. In both tasks, DG-PG converges within 10 episodes on average at every scale up to 1500 heterogeneous agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.