acceptodds
Under review as a conference paper at ICLR 2027

Faster Last-Iterate Convergence for Zero-Sum Games with Stochastic Feedback via Anchor-Regularized Learning

Abstract

We study uncoupled learning in two-player zero-sum matrix games under stochastic feedback. Each player updates from only its own feedback, without observing the opponent's action, strategy, or payoffs. We propose anchor-regularized follow-the-regularized-leader (AR-FTRL), which repeatedly solves the game regularized toward an anchor strategy profile, updating the anchor to each solution. As a warm-up, under unbiased noisy gradients with bounded second moment, the norm of the gap of the last strategy converges at rate . This guarantee exploits the metric subregularity of the zero-sum game's bilinear structure, under which the regularized equilibrium contracts toward the Nash equilibrium set by a constant factor at each anchor update. Our main result extends this to bandit feedback, where each player observes only the scalar payoff of its sampled action. Under a unique fully mixed Nash equilibrium, the norm of the gap of the last strategy converges at the same rate after observations, even though the importance-weighted gradient estimator has unbounded second moment near the boundary of the simplex. Both guarantees hold only at the anchor updates, not at every round. Under bandit feedback, this yields a faster rate than the of every-round guarantees, though the guarantee is no longer anytime.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.