Aggregate in the Advantage, Not the Ratio: A Canonical-Form Analysis of Cooperative Multi-Agent Policy Optimization
Abstract
Multi-agent policy optimization, exemplified by PPO-based methods, is a key branch of cooperative Multi-Agent Reinforcement Learning (MARL). A central design question is how many neighboring agents to aggregate so as to effectively utilize global information for cooperation. This decision must be made along two dimensions: in the advantage (which agents' rewards contribute to the credit signal) and in the ratio (which agents' likelihood ratios form the clipped importance weight). We formalize these two design choices as a pair of support patterns, one for the advantage and one for the ratio, and prove a canonical structure: the expected multi-agent policy optimization objective depends on the two patterns only through a single combined quantity, the product of the ratio pattern and the advantage pattern. This yields two consequences: (i) Redundancy: the two patterns are interchangeable with respect to the resulting learning signal, so neither way of aggregating neighbors is inherently superior to the other. (ii) Variance Ordering: the advantage combines neighbors' rewards additively, so its variance grows only mildly and it admits an interior sweet spot where the aggregation range matches the coupling neighborhood, whereas the ratio combines neighbors' likelihood ratios multiplicatively, so its variance grows exponentially with the number of aggregated agents and brings no compensating reduction in bias. Together these results point to an unambiguous design principle: aggregate neighbors in the advantage, sized to the coupling neighborhood, and keep the ratio per-agent. This also explains why neighbor-based advantages are far more common than neighbor-based ratios across prior heuristic and empirical designs. We prove these results under specified assumptions and validate them across four carefully designed synthetic cooperative games and a real-world large-scale traffic-signal control task. The code for our experiments is available in an anonymous repository at https://anonymous.4open.science/r/MAPO-803E .
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.