Improving Last-Iterate Convergence in Monte Carlo Counterfactual Regret Minimization
Abstract
Counterfactual regret minimization (CFR) and its Monte Carlo sampled variants (MCCFR) are the standard solvers for two-player zero-sum extensive-form games and the engine behind superhuman poker programs. Their guarantees cover the *average* strategy, which is cheap to track offline. A deployment that acts with the current strategy, for instance online adaptation to a changing opponent, has only the *last* iterate to act on. Under MCCFR sampling the last iterate of the regret-matching family settles at a noise floor and improves far more slowly than the average. We introduce reward-transformed MCCFR (RT-MCCFR), a regret matching plus (RM) algorithm whose last iterate lowers this floor and keeps falling from 5M to 40M node visits in our sweeps. A perturbation keeps every reach probability bounded below, an anchor regularization makes each local subproblem strongly concave, and a windowed average of the recent iterates is the output that the regularization pulls toward, at constant memory per information set. On five games under external sampling (ES) and outcome sampling (OS) with 10 seeds, the windowed anchor is below the best baseline under the same window rule in six of ten cells and above it in four, one cell resolved in each direction, and below every baseline last iterate; the raw last iterate is below every RM-family last iterate. Multiplying the budget by eight from an exploration transient lowers the anchor by a factor of 1.6 to 23 where RM moves by at most 30%, and past that budget no anchor change is resolved. On the ES noise axis the floor rises with exponent 0.47 and bootstrap half-width 0.10 in the total sampling variance, consistent with the square-root heuristic of the self-normalized RM step size, and it falls roughly inversely in the anchor strength over the working range. We prove an inverse-square-root last-iterate convergence bound, in squared distance to the regularized fixed point, in a stylized push regime.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.