CGPO: Counterfactual Credit for Multi-Agent Generative Policy Optimization
Abstract
Diffusion policies offer expressive action representations for cooperative multi-agent control, but learning from shared team rewards requires distinguishing each agent's contribution to joint performance. We introduce Counterfactual Generative Policy Optimization (CGPO), an on-policy actor-critic framework that trains decentralized diffusion policies through counterfactual credit assignment. For each agent, a centralized critic estimates a counterfactual advantage by comparing the value of its executed action with the average value of alternatives independently sampled from its behavior policy, while holding teammates' actions fixed. This advantage weights PPO-style updates at each denoising step along the diffusion trajectory that generated the executed action. Our theoretical analysis shows that estimating the counterfactual baseline with independent samples introduces an additional gradient covariance term that scales as , and identifies conditions under which counterfactual credit reduces policy-gradient variance. Experiments on VMAS, MaMuJoCo, and MPE show that CGPO achieves the highest mean returns among the compared methods, with gains of 6.7-46.6% over the strongest compared baselines on MaMuJoCo.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.