Stabilizing Independent Learning in General-Sum Self-Play via Coupling-Adaptive Regularization
Abstract
Independent policy updates in general-sum self-play offer computational efficiency but introduce severe non-stationarity. To guarantee joint objective improvement, existing theory bounds the induced independent-to-joint objective gap and constrains policy updates. The dependence of these bounds on the total agent count can make permissible steps vanish in large systems. We derive a bound on this gap in which each agent's update is weighted by its state-dependent coupling with the others. Under interaction sparsity, the coupling coefficient for each agent remains as the population grows. Building on this analysis, we introduce Coupling-Adaptive Regularized Policy Optimization (CARPO). CARPO uses a vector critic to estimate local coupling and adapts each agent's trust region, tightening updates in highly coupled states and relaxing them in loosely coupled ones. Empirical evaluations range from a tabular game to large-scale autonomous driving. By restricting updates during critical interactions, CARPO induces a “safe hesitation” phase that prevents premature convergence in a coordination trap. Its state-dependent adaptation outperforms the tested static clipping choices across coordination levels. In autonomous driving, CARPO maintains the rapid initial learning of independent updates while surpassing the final performance of the joint-update baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.