Minimax-Optimal Robust Reinforcement Learning with -divergence Uncertainty Set under - and -Rectangularity
Abstract
We study the sample complexity of model-based distributionally robust reinforcement learning (DR-RL) with -divergence defined uncertainty set under balanced nominal generative-model sampling. For -rectangular uncertainty set, existing upper and lower bounds do not match in the uncertainty level . We close this gap by establishing a tighter upper bound of , which matches with the minimax lower bound and thus we establish the minimax optimality. Here, is the discount factor, and are the state and action spaces, and is the optimality gap. Our proof develops a novel variance-sensitive framework that exploits the second-moment structure of adversarial density ratios to control robust Bellman estimation and to propagate statistical errors while preserving dependence linear in the divergence. For -rectangular uncertainty set with shared state-wise budget of , we further establish an upper bound of and also a matching minimax lower bound, and therefore we establish the minimax optimality also for -rectangular uncertainty set. Our results show that the sample complexity shall not depend on the minimum of non-zero transition probability of the nominal transition kernel, and provide the correct scaling with the uncertainty level .
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.