Belief-Based Risk Reduction for Robust Reinforcement Learning under Adversarial Observation Perturbations
Abstract
Deep reinforcement learning (DRL) policies are highly vulnerable to adversarial noise in observations, posing significant risks in safety-critical scenarios. The primary challenge of adversarial perturbations lies in their ability to distort the information observed by an agent. Recently, state purification defenses have emerged as a promising solution. However, existing methods suffer from a distribution mismatch between clean states and adversarial observations, as well as uncertainty in state restoration and risks in decision-making. To address these challenges, we propose Belief-Based Risk Reduction (BBRR), a risk-aware state purification framework. BBRR uses a diffusion bridge to mitigate the distribution gap between adversarial observations and clean samples, while employing belief state approximation and Conditional Value-at-Risk (CVaR) to reduce restoration uncertainty and decision risks, respectively. Theoretical analysis and comprehensive evaluations across diverse environments demonstrate that BBRR consistently achieves state-of-the-art performance against a range of strong attacks, validating its robustness and practical applicability. Our code is available at https://anonymous.4open.science/r/Robust_BBRR.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.