acceptodds
Under review as a conference paper at ICLR 2027

Expert-Regularized Equilibrium Computation via Convex Markov Games

Abstract

We investigate the optimization landscape of Markov game equilibria that are anchored on (expert) reference policies. The technique of anchoring, coupled with self-play training, has achieved super-human performance in several games. Despite its strong empirical success, its formal analysis has lagged behind and has so far been restricted to stateless matrix games and/or direct policy parametrizations. We show that any reference-regularized Markov game can naturally be placed within the framework of a convex Markov game (cMG). This introduces a novel instance of “imitation” within the cMG framework—complementing previously recognized aspects such as behavioral imitation, creativity, safety, and equilibrium fairness. Framing the setting in this way, we establish two complementary convergence results. First, in the unconstrained case, we prove that simultaneous gradient descent-ascent converges locally to a (regularized) Nash equilibrium, offering theoretical support for the simultaneous-update schemes that dominate practice. Such guarantees, however, cannot in general be promoted to global ones: simultaneous GDA is known to fail globally even under a two-sided Polyak-Łojasiewicz condition. To obtain global convergence, we turn to softmax policies under alternating updates. We prove that saddle-point computation satisfies a two-sided (proximal) Polyak-Łojasiewicz condition for softmax parametrizations and design an alternating policy gradient scheme that converges to a (regularized) Nash equilibrium with an explicit polynomial rate and sample complexity. In doing so, we obtain the first convergence proof of softmax policy gradient to a Nash equilibrium in convex Markov games, offering a crisp theoretical justification for the use of softmax policy gradient updates in numerous practical applications.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.