Certified Diffusion Smoothing for Robust Reinforcement Learning
Abstract
Deep Reinforcement Learning (DRL) is known to be vulnerable to small adversarial perturbations in the input space, which limits its deployment in real-world applications. While many empirical studies have sought to improve its robustness against adversarial attacks, these methods typically lack formal robustness guarantees. Existing certified defenses, despite providing certification criteria, often suffer from low robust rewards and weak robustness guarantees. In this work, we introduce a certifiably robust reinforcement learning framework that leverages the diffusion process to incorporate the objective of randomized smoothing directly into policy training. We propose Certified Diffusion Smoothing (CDS), a novel method that trains a smoothed agent with stronger certified robustness guarantees and higher robust rewards. We evaluate CDS under multiple certifications and adversarial attacks in standard RL environments, including MuJoCo and Atari. Across these benchmarks, CDS consistently outperforms previous smoothed agents by an average factor of 1.5–3.2 and robustly trained agents by an average factor of 2.8. Our code is available at https://anonymous.4open.science/r/CDS_DRL.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.