Span-Based Analysis of Distributionally Robust Average-Reward Markov Games
Abstract
We study finite distributionally robust Markov games under the average-reward criterion, where each player evaluates the worst-case average-reward of policies against a player-specific, state–joint-action rectangular ambiguity set. We introduce the uniform optimal-response bias span and develop studies on both equilibrium existence and learning guarantees in terms of this single parameter. Firstly, we show under different chain structures and uncertainty sets, and that a finite guarantees the existence of a stationary statewise robust average-reward Nash equilibrium. We then establish the convergence of Nash equilibria under discounted reward to average reward as the discount factor approaching , and develop a reduction framework to connect them. Furthermore, under total-variation ambiguity and balanced nominal generative-model sampling, we show that a model-based discounted learner with any exact discounted Nash oracle uses samples to learn an -robust average-reward Nash equilibrium, where are the state and joint action space sizes, and is the largest next-state support size. This guarantee holds simultaneously for every exact empirical Nash oracle. A matching lower bound on two-player, two-action games further proves the support factor unavoidable for this balanced, every-oracle interface, implying the tightness of our upper bound.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.