acceptodds
Under review as a conference paper at ICLR 2027

VPPO: Verified Pathwise Policy Optimization

Abstract

Proximal policy optimization (PPO) reuses on-policy data through a clipped surrogate, but clipping alone does not explicitly bound policy change. Incorporating critic action gradients offers a complementary direction for policy improvement, provided that their reliability can be assessed. We introduce Verified Pathwise Policy Optimization (VPPO), an on-policy algorithm that combines clipped natural-gradient updates with probe-calibrated pathwise refinement. More specifically, VPPO performs multiple re-linearized steps on each rollout batch, subject to explicit KL bounds and nondecrease of the regularized surrogate. To calibrate trust in the critic-guided refinement, we use paired simulator trajectories with common random numbers to evaluate the proposed action direction. A damped-Fisher projection then preserves first-order compatibility between the refinement and surrogate directions. A matched-budget ablation shows that clipping remains necessary for stable batch reuse even under explicit KL control. On GPU-parallelized benchmarks, VPPO improves humanoid locomotion performance over the standard PPO configuration while achieving comparable final returns on additional control tasks. Under matched entropy regularization, VPPO reaches intermediate performance milestones earlier than PPO. Compared with off-policy baselines, VPPO also reaches target returns in less wall-clock time on most tested control tasks, despite equiring more environment steps.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.