ESPO: Early-Stopping Proximal Policy Optimization
Abstract
When a reasoning trajectory becomes unrecoverable, standard reinforcement learning algorithms still generate until an end-of-sequence (EOS) token is generated or the rollout horizon is reached. This can waste computation and add estimation noise from post-failure tokens. We propose (arly-topping Proximal olicy ptimization), which identifies likely trajectory failures on-the-fly and terminates rollouts early. At each generation step, ESPO computes a surrogate regret using only the logits already computed during sampling, and terminates when the normalized, smoothed deviation exceeds a critic-dependent threshold. Truncated trajectories are treated as absorbing failure states with a terminal penalty at the detected stopping step, propagated through the retained prefix by GAE, without an additional reward model or human annotation. Across three backbones from two model series, on both mathematical reasoning and code generation, ESPO reaches on AIME 2024, on AMC 2023 and on MATH-500 with DeepSeek-R1-Distill-Qwen-7B — each the best figure among the five methods we compare, as is its average over six benchmarks — while decoding 21.7% fewer rollout tokens than PPO and cutting per-step wall-clock by 23.2%. Our code is available in the supplementary material.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.