GIRO: Robust Reinforcement Learning with Generative Model Induced Uncertainty Set
Abstract
Offline model-based reinforcement learning uses learned transition dynamics to generate training experience, but policies optimized against an estimated model may exploit its errors and fail under distribution shift. We introduce GIRO, a distributionally robust RL method that defines an uncertainty set of generative transition models through a training-loss budget. Because shared model parameters couple transitions across state–action pairs, the resulting set is generally non-rectangular, making robust policy evaluation challenging. GIRO alternates adversarial model updates with robust policy improvement. For robust policy evaluation, it uses a primal-only update that decreases policy value when the loss constraint is satisfied and reduces constraint violation otherwise. Under stated assumptions, we bound robust evaluation suboptimality and constraint violation and establish a stationarity guarantee for policy improvement. Across nine D4RL task–dataset combinations and eight test perturbations, GIRO attains the highest aggregate worst-case return and lowest aggregate across-perturbation variance among the compared methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.