acceptodds
Under review as a conference paper at ICLR 2027

Last-Iterate Convergence of Parameter-Free and Scale-Invariant Regret Matching

Abstract

Regret Matching (RM) is widely used as a local regret minimizer in counterfactual regret minimization (CFR) for two-player zero-sum extensive-form games. RM and its predictive variants are parameter-free and invariant to positive rescaling of the payoffs, and existing guarantees control their average or best iterates. Their last-iterate behavior, however, has remained theoretically unresolved even in two-player zero-sum matrix games, where their self-play dynamics can be studied in isolation. We study two predictive RM dynamics in this setting. The first is the Increasing Regret Extra-Gradient Predictive Regret Matching (IREG-PRM) dynamics of zhang2026, and the second is the Increasing-Regret Optimistic PRM (IRO-PRM) dynamics that we introduce. We prove that the played strategy profiles of both dynamics converge to the Nash equilibrium set for every two-player zero-sum matrix game. For IREG-PRM, we further bound the cumulative duality gap by . To the best of our knowledge, these are the first last-iterate convergence guarantees for a parameter-free and payoff-scale-invariant RM-type algorithm. Experiments on structured and random matrix games support the theoretical results.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.