Constant Swap Regret in General-Sum Games via Two-Scale Higher-Order Optimism
Abstract
We give deterministic and uncoupled learning dynamics for finite multiplayer general-sum games under full-information feedback that achieve constant individual swap regret in self-play, independent of the horizon . With players and at most actions each, every player’s individual swap regret is at every finite horizon. The dynamics use the classical Blum–Mansour framework with optimism. Each player predicts the deviation gains, uses these predictions to update a row-stochastic transition matrix, and plays its stationary distribution. Our new ingredients include a tailored row normalization map and a two-scale higher-order predictor. An adversarially robust variant, obtained through a generic common-prefix switching wrapper, preserves the self-play bound up to a universal constant and guarantees individual swap regret at most in the adversarial setting.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.