When Does Local Credit Assignment Improve Shared Policy Gradients in Multi-Agent Reinforcement Learning?
Abstract
Credit assignment in multi-agent reinforcement learning is justified agent by agent: a counterfactual or teammate-marginalized critic can reduce the variance of a policy-gradient term. With parameter sharing, the optimizer updates the sum of agent-time contributions, so local variance reduction need not improve the shared gradient. This gap is fundamental. For every team size, we construct bounded games with an exact critic and a nonzero shared gradient in which every locally variance-optimal correction reduces local variance but increases shared-gradient variance by an arbitrarily large factor. For two agents sharing a Bernoulli policy with the value baseline, we identify a sharp regime in which exact teammate marginalization is guaranteed not to increase shared-gradient covariance, and extend it to an exact criterion for identical agents. These results motivate PLACID, which separates admission - whether and at what strength a correction improves the shared-gradient estimate - from realization - how to compute it with less Monte Carlo noise. On 16 fixed tabular games, empirical admission removes up to 9.21% of sparse population mean-squared error. On the two fitted critic families, it exceeds unit marginalization at every tested batch budget. Its certified variant selects 30,101 corrections with none increasing population shared-gradient variance. In 1,566 enumerated games, even population-moment local screens can select corrections that harm the shared gradient. In shared-actor training, a lagged empirical admission variant increases normalized return AUC on coverage games by over the sparse update and exceeds unit marginalization on team and coverage games with one teammate proposal.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.