acceptodds
Under review as a conference paper at ICLR 2027

Distributionally Robust Average-Reward Reinforcement Learning: Finite-Sample Guarantees under Weakly Communicating

Abstract

We study distributionally robust reinforcement learning (DR-RL) in the average-reward setting under weak communication. Our main result provides finite-sample guarantees for estimating the robust optimal average reward and learning a near-optimal policy, covering both SA-rectangular and S-rectangular structures with divergence-based and distance-based uncertainty sets. Specifically, for Kullback–Leibler and -divergence balls, we establish explicit radius conditions under which the robust average-reward Bellman equation admits a constant-gain solution, while for total variation and Wasserstein balls, any positive radius suffices without requiring the nominal MDP is weakly communicating. Our algorithm is prior-knowledge-free and achieves sample complexities of for estimating the robust optimal average reward and for learning an -optimal policy. Here, is the smallest positive nominal transition probability and is a robust optimal bias function. We further provide an almost-tight explicit upper bound on . Finally, we validate the predicted convergence rate through numerical experiments.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.