acceptodds
Under review as a conference paper at ICLR 2027

Leveraging Generative Uncertainty Modeling for Distributionally Robust Offline Reinforcement Learning

Abstract

Distributionally robust offline RL addresses the challenge of distribution shift of deployment environment by optimizing policies against an uncertainty set of environment dynamics. The uncertainty-set design is critical: it should capture diverse distribution shifts for robustness while excluding unrealistic shifts that lead to overly pessimistic policies. Existing methods often assume rectangular uncertainty, which can combine locally worst-case transitions that cannot coexist in a single environment. Many also rely on standard distribution discrepancy measures: -divergences ignore support-space geometry, while Wasserstein sets can be computationally challenging in high-dimensional continuous settings. To capture non-rectangular, diverse yet plausible shifts, we construct a generative uncertainty set using flow models which can generate diverse and plausible trajectories. The proposed set captures temporally coupled dynamics shifts through shared model parameters while maintaining consistency with offline data via a tractable flow-matching constraint that implies a Wasserstein bound. Building on this uncertainty model, we develop Generative Uncertainty-Augmented Robust Offline RL (GURO-RL), which addresses value evaluation and adversarial optimization under flow-based uncertainty and finite offline data. Experiments across multiple applications demonstrate that GURO-RL improves policy robustness under diverse distribution shifts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.