acceptodds
Under review as a conference paper at ICLR 2027

ISPO: Correcting Policy-Induced Value Lag in PPO via Importance-Sampled Policy Shifts

Abstract

On-policy actor–critic RL, represented by PPO, trains the critic on rollouts from the policy it just replaced. After each actor update, the true value function moves while the critic still fits the previous one, leaving a **policy-induced value lag**. Our analysis shows that this lag biases every advantage estimate the actor consumes and, when large, violates the small-update assumption that PPO relies on. The lag is bounded by the product of effective horizon, action-wise residual dispersion, and policy-update size, identifying long-horizon, action-sensitive control as its worst case. Prior work addresses this through trajectory-level reweighting (Retrace, V-trace), which compounds importance-ratio variance, or policy-conditioned values (PeVFA), which add parameters and training costs. We argue neither is necessary: PPO constrains policy updates, making the value shift nearly first-order and estimable directly from the existing rollout. Motivated by this, we propose **I**mportance-**S**ampled **P**olicy-Shift **O**ptimization (**ISPO**). We construct a **PINN-style evaluation residual** per transition based on the policy evaluation PDE, and reweight it with a **single-step importance ratio**, avoiding cumulative off-policy weights on the whole trajectory. The resulting shift corrects the critic target after each actor update toward the updated policy's value function. We evaluate ISPO on an analytic LQR system, standard MuJoCo benchmarks, and an embodied manipulation task. Experiments show that ISPO achieves higher final return while adding no environment interaction, no learnable parameters, and no cumulative importance weights, improving return by 10% to 40% over vanilla PPO and matching or exceeding other baselines. [Here is Project Page](https://anonymous.4open.science/r/ISPO-FFF6/)

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.