acceptodds
Under review as a conference paper at ICLR 2027

Coupled Dual-Adversarial Regularized Self-Play

Abstract

Naive self-play for two-player zero-sum games can cycle in nontransitive games. Regularized self-play stabilizes training with a single network but only plays against itself, so it forgets earlier opponents and cannot use external ones; population methods keep an opponent pool, but their equilibrium lives in a mixture over the pool, and heuristic weightings such as PFSP carry no convergence guarantee. We propose Coupled Dual-Adversarial Regularized self-play (CoDA-R). Each round, CoDA-R sets the pool weights by a Euclidean projection anchored at pure self-play, so that only opponents currently beating the latest policy receive weight, and then takes a proximal policy step against the resulting mixture. The coupling yields one energy inequality that charges every policy step and every unit of pool weight to the remaining distance to a Nash equilibrium. For the exact algorithm, this gives last-iterate convergence of the latest policy to a Nash equilibrium however oracle opponents enter or leave the pool, plus certificates computable during training and error-tolerant versions of all bounds. In OpenSpiel at a matched budget, the latest network of CoDA-R cuts the tail exploitability on Goofspiel(4) from 0.797 for R-NaD-PPO to 0.083, on par with a multi-network PSRO mixture; on Leduc poker, with the dual strength scaled to the payoff range and an exact best-response oracle, it roughly halves that of R-NaD-PPO; on Liar's Dice the two tie.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.