Same Noise, Better Gradients: Policy Optimization with Common Random Numbers
Abstract
We study common random numbers as a variance-reduction method for policy- gradient reinforcement learning in stochastic simulators. The core idea is to decouple the random number generators for action selection and state transitions, and investigate counterfactual effects of actions under the same state transition noise. For policy gradient algorithms, we implement this idea as a control-variate that depends on the state transition noise, and prove that it is unbiased and obtains lower variance in certain conditions. Further, we propose two simple modifications of the popular PPO algorithm based on either (1) resetting the random number generator, and (2) learning a value function that depends on the random seed. We demonstrate improved performance on various stochastic domains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.