acceptodds
Under review as a conference paper at ICLR 2027

Convergence and Implicit Bias of Softmax Extragradient Flow in Bilinear Zero-Sum Games

Abstract

Many learning problems optimize over probability distributions represented by softmax logits. Although extragradient enjoys strong guarantees for convex-concave games in policy space, its behavior under softmax parameterization is less understood: even bilinear games become nonconvex-nonconcave in logit space, and boundary equilibria correspond to logits at infinity. We study the high-resolution ODE of the extragradient method in softmax-parametrized skew-symmetric bilinear zero-sum games under symmetric self-play. We establish ***global convergence*** under two structural conditions. First, if there is an interior equilibrium, every trajectory converges in logit space. Second, under a condition we call *flat dominance*, we show convergence of every trajectory to the Nash set, with explicit rates. We further characterize the initialization-dependent ***implicit bias*** of the dynamics: which policy is selected? Conditional on policy convergence, we show that the trajectory must converge to a Nash equilibrium (NE). Under flat dominance, almost every such initialization leads to a *pure-strategy* NE. We further prove policy convergence for two structured classes of games and derive an explicit initialization-dependent NE selection rule. Finally, we present numerical experiments illustrating our theoretical results.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.