acceptodds
Under review as a conference paper at ICLR 2027

Directional-Clamp PPO

Abstract

Proximal Policy Optimization (PPO) is one of the most widely used algorithms in deep reinforcement learning. Its objective encourages the importance ratio between the current and behavior policies to move in the "right direction”, i.e., increasing the ratios for positive advantages and decreasing them for negative ones. To prevent overly aggressive updates in the right direction, PPO and its variants employ clipping mechanisms that restrict the magnitude of such updates. However, due to the stochastic nature of the training process, the importance ratios frequently move in the "wrong direction” during PPO optimization. We identify the strictly wrong direction updates as a key but largely overlooked factor that limits the performance of PPO. To address this, we propose the Directional-Clamp PPO algorithm (DC-PPO), which further penalizes the actions going to the "strict wrong direction” regions, where the advantage is positive (negative) and importance ratio falls below (above) (), for a tunable parameter . The penalty is by enforcing a steeper loss slope, i.e., a clamp, in those regions. We demonstrate across a range of environments that DC-PPO consistently outperforms PPO and its variants that modify the objective's behavior in the right direction. Moreover, we show, both theoretically and empirically, that DC-PPO better avoids strict wrong direction updates while keeping the importance ratio closer to . Finally, we prove global convergence for tabular DC-PPO with direct policy parameterization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.