Grouped Heterogeneous-Agent Proximal Policy Optimization
Abstract
Proximal Policy Optimization (PPO) has been successfully extended to cooperative multi-agent reinforcement learning (MARL), providing an effective framework for learning coordinated policies. Multi-Agent Proximal Policy Optimization (MAPPO) achieves strong sample and parameter efficiency through parameter sharing, but a fully shared policy can induce homogeneous behaviors and fail to represent asymmetric optimal solutions. Heterogeneous-Agent Proximal Policy Optimization (HAPPO) supports heterogeneous behaviors by maintaining separately parameterized policies and coordinating agent-by-agent sequential updates, but the number of dependent update stages grows with the number of agents, limiting training efficiency. In this paper, we propose Grouped Heterogeneous-Agent Proximal Policy Optimization (GHAPPO). By extending multi-agent advantage decomposition to groups, GHAPPO shares a policy backbone within each group and updates groups sequentially, thereby reducing the number of dependent update stages. In addition, we introduce a lightweight agent-specific action-replacement module to promote policy heterogeneity within each group. Under an explicit bounded-ratio assumption, our analysis characterizes the theoretical relationship between grouping granularity and sequential-update overhead, while an example shows that action replacement can raise the attainable performance upper bound relative to full parameter sharing. Experiments on SMAC and SMACv2 demonstrate that GHAPPO improves training efficiency while maintaining strong performance across different scenarios.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.