acceptodds
Under review as a conference paper at ICLR 2027

ESPO: Early-Stopping Proximal Policy Optimization

Abstract

When a reasoning trajectory becomes unrecoverable, standard reinforcement learning algorithms still generate until an end-of-sequence (EOS) token is generated or the rollout horizon is reached. This can waste computation and add estimation noise from post-failure tokens. We propose (arly-topping Proximal olicy ptimization), which identifies likely trajectory failures on-the-fly and terminates rollouts early. At each generation step, ESPO computes a surrogate regret using only the logits already computed during sampling, and terminates when the normalized, smoothed deviation exceeds a critic-dependent threshold. Truncated trajectories are treated as absorbing failure states with a terminal penalty at the detected stopping step, propagated through the retained prefix by GAE, without an additional reward model or human annotation. Across three backbones from two model series, on both mathematical reasoning and code generation, ESPO reaches on AIME 2024, on AMC 2023 and on MATH-500 with DeepSeek-R1-Distill-Qwen-7B — each the best figure among the five methods we compare, as is its average over six benchmarks — while decoding 21.7% fewer rollout tokens than PPO and cutting per-step wall-clock by 23.2%. Our code is available in the supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.