Distributionally Robust Average-Reward Reinforcement Learning: Finite-Sample Guarantees under Weakly Communicating
Abstract
We study distributionally robust reinforcement learning (DR-RL) in the average-reward setting under weak communication. Our main result provides finite-sample guarantees for estimating the robust optimal average reward and learning a near-optimal policy, covering both SA-rectangular and S-rectangular structures with divergence-based and distance-based uncertainty sets. Specifically, for Kullback–Leibler and -divergence balls, we establish explicit radius conditions under which the robust average-reward Bellman equation admits a constant-gain solution, while for total variation and Wasserstein balls, any positive radius suffices without requiring the nominal MDP is weakly communicating. Our algorithm is prior-knowledge-free and achieves sample complexities of for estimating the robust optimal average reward and for learning an -optimal policy. Here, is the smallest positive nominal transition probability and is a robust optimal bias function. We further provide an almost-tight explicit upper bound on . Finally, we validate the predicted convergence rate through numerical experiments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.