Capacity Scaling Fragility in Multi-Agent Policy Optimization
Abstract
Increasing network capacity is generally expected to improve the performance of reinforcement-learning policies, or at least not to harm it. However, we find that this expectation can fail in cooperative multi-agent reinforcement learning (MARL), because network width interacts with how each policy network is up- dated. On SMACv2 protoss 5 vs 5, under standard tuned hyperparameters, in- creasing HAPPO’s hidden size from 64 to 256 reduces the win rate from 0.474 to 0.368 over five seeds, whereas MAPPO does not show the same degradation. A four-way identification experiment crossing sequential versus synchronous up- dates with independent versus shared parameters shows that the capacity penalty appears in every standard configuration except shared, synchronous MAPPO. Further interventions transfer this protection to independent actors, showing that parameter sharing itself cannot explain the difference. Across these configura- tions and interventions, the key quantity that consistently separates protected and fragile regimes is the realized policy-update scale. At the same nominal learn- ing rate, increasing network width amplifies the net per-iteration policy KL by 2.2–3.6×, while fragile topologies exhibit 2.4–12× larger absolute policy dis- placement than protected ones. When the larger network’s realized KL is brought back to the small network’s default level, its performance is restored. Reduc- ing the learning rate shrinks the capacity gap in every fragile configuration we test, while directly bounding per-update KL provides a partial recovery. Together, these results suggest that network capacity, update topology, and learning rate jointly determine the realized optimization regime: the same nominal learning rate can produce substantially different policy updates across network widths and training configurations. At hidden size 256, the default-rate comparison gives MAPPO a 0.094 performance advantage over HAPPO; applying the same lower learning rate to both algorithms reduces this margin roughly fourfold to 0.022. Thus, fixed-architecture, fixed-learning-rate MARL comparisons can confound algorithmic differences with update-scale calibration.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.