Which Nash Equilibrium? Solver-Dependent Selection on Zero-Sum Nash Polytopes
Abstract
Regularized self-play—the algorithm family behind DeepNash's master-level Stratego play—drives a two-player zero-sum policy to a Nash equilibrium by repeatedly best-responding to an entropy-regularized, slowly moving reference policy ρ. When a game has a polytope of value-equivalent equilibria, this regularizer silently breaks the tie, and with a uniform reference it selects the maximum-entropy member. We ask whether the reference can be used to pick the equilibrium on purpose. We study this on five games plus a two-dimensional Nash polytope, where the full Nash set and exact best responses are known, and back every claim with seed-bootstrap confidence intervals, equivalence tests, or hypothesis tests. Anchoring the reference at a target member and running ordinary refinement steers self-play to that member with mean coordinate error 0.007 (95% CI [0.002, 0.015]) at median exploitability 5e-5 (TOST equivalence within ±0.05, p=3e-16). The anchoring persists through refinement, and the landing point follows the reference, not the initialization. The selected member approximately follows the reach-weighted information projection of ρ onto the polytope (slope 0.969 [0.950, 0.987], R²=0.993). We report the failure modes just as prominently. A fixed off-manifold reference steers only at an exploitability cost of 0.08–0.25. Too large a mirror step makes runs look converged while failing to express the selection. Targets near the boundary of the stiffest family undershoot. Table and MLP steering maps are statistically equivalent; attention adds seed variance. Finally, against a best response the selection–robustness trade-off is degenerate. The recipe—initialize the reference at the desired member and refine—recasts the KL-to-reference term of trust-region and RLHF-style RL as a selection knob, not only a stabilizer.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.