acceptodds
Under review as a conference paper at ICLR 2027

Revisiting Amari's -Divergence to improve clipped surrogate PPO

Abstract

Proximal Policy Optimization (PPO) is one of the most popular policy gradient algorithms due to its widespread practicality and scalability. However, its performance is known to be highly sensitive to increasing the clipping threshold and policy network depth. To address this sensitivity, we introduce ALDIPPO, which replaces the clip with an -divergence penalty on the importance ratio. This makes the penalty stiffer for samples outside the range, which PPO rejects by switching their gradient off, wasting the rollout budget spent to collect them. We devise several online update rules for based on on-policy batch statistics, and show that ALDIPPO beats PPO from toy MiniGrid environments to standard MuJoCo tasks, and generalizes better in the Procgen benchmark.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.