One Gradient, Yet More Robust Policies: Momentum-Guided Policy Optimization
Abstract
Robust reinforcement learning (Robust RL) aims to learn policies that remain effective under environmental uncertainty. Recent methods pursue such robustness by seeking flat reward maxima, often through sharpness-aware minimization (SAM), which encourages high returns under small parameter perturbations. However, vanilla SAM first constructs an adversarial parameter perturbation from a stochastic policy gradient and then computes a second gradient at the perturbed parameters to update the policy. This procedure relies on noisy gradient estimates and incurs additional computation, which can make it complex and unstable. In this paper, we introduce a policy optimization method, based on Momentum-SAM (mSAM), for Robust RL that utilizes accumulated momentum to construct adversarial perturbations. The resulting method is effective and computationally lightweight, requiring only one policy-gradient evaluation per update while reducing reliance on a single noisy estimate. Furthermore, we establish finite-time stationarity bounds for SAM and mSAM in policy optimization, providing a theoretical foundation for sharpness-aware policy learning. Our experiments show the effectiveness of mSAM in enhancing the robustness of RL compared with existing methods under diverse perturbations, including actions, transition dynamics, and reward functions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.