VPPO: Verified Pathwise Policy Optimization
Abstract
Proximal policy optimization (PPO) reuses on-policy data through a clipped surrogate, but clipping alone does not explicitly bound policy change. Incorporating critic action gradients offers a complementary direction for policy improvement, provided that their reliability can be assessed. We introduce Verified Pathwise Policy Optimization (VPPO), an on-policy algorithm that combines clipped natural-gradient updates with probe-calibrated pathwise refinement. More specifically, VPPO performs multiple re-linearized steps on each rollout batch, subject to explicit KL bounds and nondecrease of the regularized surrogate. To calibrate trust in the critic-guided refinement, we use paired simulator trajectories with common random numbers to evaluate the proposed action direction. A damped-Fisher projection then preserves first-order compatibility between the refinement and surrogate directions. A matched-budget ablation shows that clipping remains necessary for stable batch reuse even under explicit KL control. On GPU-parallelized benchmarks, VPPO improves humanoid locomotion performance over the standard PPO configuration while achieving comparable final returns on additional control tasks. Under matched entropy regularization, VPPO reaches intermediate performance milestones earlier than PPO. Compared with off-policy baselines, VPPO also reaches target returns in less wall-clock time on most tested control tasks, despite equiring more environment steps.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.