Bayesian Adversarial Distributional Reinforcement Learning for Sample-Efficient Risk-Sensitive Control
Abstract
Distributional reinforcement learning (DRL) enables risk-sensitive decision-making by learning the distribution of cumulative returns rather than only their expectation. However, accurately estimating return distributions can require substantial interaction with the environment, making DRL difficult to apply when data collection is risky, costly, or limited. With limited experience, uncertainty in the learned return distribution can remain substantial, yet many widely used DRL methods do not explicitly quantify epistemic uncertainty in the learned return model. We propose GAN-DRL, a Bayesian Generative Adversarial Network (GAN) framework for distributional reinforcement learning that combines Wasserstein-based return-distribution learning with Bayesian inference to quantify epistemic uncertainty in the learned return model. We use posterior information gain as an intrinsic signal to guide the agent toward informative experience, improving learning efficiency in stochastic risk-sensitive environments. We evaluate GAN-DRL on two continuous-control problems: European option hedging and F1TENTH autonomous driving. Across both domains, GAN-DRL reaches comparable or improved risk-sensitive performance using substantially fewer environment interactions than distributional and actor-critic baselines. In European option hedging, GAN-DRL improves tail-risk performance, as measured by conditional value-at-risk (CVaR), while maintaining competitive average performance. In autonomous driving, GAN-DRL improves lower-tail returns and reduces rollover rates on both the Offroad-Flat and Offroad-Bumpy tasks. Overall, these results show that uncertainty-guided distributional learning can substantially improve sample efficiency while maintaining or improving risk-sensitive performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.