Directional-Clamp PPO
Abstract
Proximal Policy Optimization (PPO) is one of the most widely used algorithms in deep reinforcement learning. Its objective encourages the importance ratio between the current and behavior policies to move in the "right direction”, i.e., increasing the ratios for positive advantages and decreasing them for negative ones. To prevent overly aggressive updates in the right direction, PPO and its variants employ clipping mechanisms that restrict the magnitude of such updates. However, due to the stochastic nature of the training process, the importance ratios frequently move in the "wrong direction” during PPO optimization. We identify the strictly wrong direction updates as a key but largely overlooked factor that limits the performance of PPO. To address this, we propose the Directional-Clamp PPO algorithm (DC-PPO), which further penalizes the actions going to the "strict wrong direction” regions, where the advantage is positive (negative) and importance ratio falls below (above) (), for a tunable parameter . The penalty is by enforcing a steeper loss slope, i.e., a clamp, in those regions. We demonstrate across a range of environments that DC-PPO consistently outperforms PPO and its variants that modify the objective's behavior in the right direction. Moreover, we show, both theoretically and empirically, that DC-PPO better avoids strict wrong direction updates while keeping the importance ratio closer to . Finally, we prove global convergence for tabular DC-PPO with direct policy parameterization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.