Revisiting Amari's -Divergence to improve clipped surrogate PPO
Abstract
Proximal Policy Optimization (PPO) is one of the most popular policy gradient algorithms due to its widespread practicality and scalability. However, its performance is known to be highly sensitive to increasing the clipping threshold and policy network depth. To address this sensitivity, we introduce ALDIPPO, which replaces the clip with an -divergence penalty on the importance ratio. This makes the penalty stiffer for samples outside the range, which PPO rejects by switching their gradient off, wasting the rollout budget spent to collect them. We devise several online update rules for based on on-policy batch statistics, and show that ALDIPPO beats PPO from toy MiniGrid environments to standard MuJoCo tasks, and generalizes better in the Procgen benchmark.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.