acceptodds
Under review as a conference paper at ICLR 2027

CGPO: Counterfactual Credit for Multi-Agent Generative Policy Optimization

Abstract

Diffusion policies offer expressive action representations for cooperative multi-agent control, but learning from shared team rewards requires distinguishing each agent's contribution to joint performance. We introduce Counterfactual Generative Policy Optimization (CGPO), an on-policy actor-critic framework that trains decentralized diffusion policies through counterfactual credit assignment. For each agent, a centralized critic estimates a counterfactual advantage by comparing the value of its executed action with the average value of alternatives independently sampled from its behavior policy, while holding teammates' actions fixed. This advantage weights PPO-style updates at each denoising step along the diffusion trajectory that generated the executed action. Our theoretical analysis shows that estimating the counterfactual baseline with independent samples introduces an additional gradient covariance term that scales as , and identifies conditions under which counterfactual credit reduces policy-gradient variance. Experiments on VMAS, MaMuJoCo, and MPE show that CGPO achieves the highest mean returns among the compared methods, with gains of 6.7-46.6% over the strongest compared baselines on MaMuJoCo.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.