Best Response Convergence
Abstract
As players update their strategies through repeated interaction within a multi-player game, they may increasingly favor one action while another is currently more profitable. Even when their behavior converges, small changes in the opponents' play can keep changing which of their actions is best. We study this distinction for follow-the-regularized-leader (FTRL). More specifically, we ask whether a fixed action eventually remains a current best response. To do this, we analyze learning near boundary equilibria—a subclass of weak Nash equilibria—where one player's limiting action is tied with several unused actions, which we call alternatives. We derive tight rates for the probabilities of these alternatives and identify a sharp best-response threshold for the Hedge algorithm, the entropy-regularized form of FTRL. With symmetric learning rates , the threshold is , independently of the number of alternatives. We also show an interesting consequence of our analysis: larger learning rates can make the gains from deviation vanish more slowly, despite faster concentration on the limiting action. Lastly, we also prove a best-response stability result for Euclidean FTRL and verify our findings using simulations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.