Off-Policy Policy Optimization
Abstract
Policy gradient methods such as proximal policy optimization (PPO) are typically trained from fresh on-policy rollouts. At the same time, strong teacher policies can provide high-quality demonstrations that expose the student to successful actions and trajectories it may rarely discover on its own. Directly incorporating such demonstrations into policy-gradient training, however, creates substantial policy and state-distribution mismatch, raising concerns about biased or unstable updates and whether learning signal from teacher trajectories remains useful for the student policy. Surprisingly, we find that PPO can learn from teacher demonstrations under extreme policy mismatch when teacher and student samples enter the same standard clipped PPO surrogate with token-level importance ratios. Under explicit assumptions on an unclipped population surrogate, we characterize how teacher demonstrations induce selective imitation, providing action guidance at student-visited prefixes and exploration guidance at prefixes the student has not yet reached, while still permitting improvement beyond the teacher. Empirically, off-policy demonstrations speed up PPO convergence by even when teacher-action importance ratios start as low as , and achieve the highest greedy test accuracy on three language tasks among supervised fine-tuning, on-policy distillation, and on-policy PPO.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.