acceptodds
Under review as a conference paper at ICLR 2027

Learning with a Single Rollout via Pass@k Policy Optimization

Abstract

Group Relative Policy Optimization (GRPO) typically requires multiple rollouts per prompt to estimate relative credit, incurring substantial sampling cost and exacerbating training asynchrony. This limitation is particularly severe for scientific reasoning agents, where each trajectory may involve long-horizon reasoning, multiple turns, and parallel sub-runs. We study single-rollout reinforcement learning, where only one rollout is collected per prompt, and find that naive training can stagnate or become severely unstable. We introduce Single-Rollout Pass@ Policy Optimization (SR-PPO), which combines Monte Carlo token-level advantage estimation with Pass@ gradient reweighting. An online critic without large-scale pretraining is used both as a prefix-dependent baseline for credit assignment and to estimate the Pass@ gradient weight. Using only one rollout per prompt, SR-PPO achieves stable learning and performance comparable to GRPO across reasoning and logic benchmarks, even under off-policy drift that degrades the performance of standard Pass@1 PPO. A key factor in this stability is Pass@ reweighting, which emphasizes harder prompts while downweighting high-success prompts, for which imperfect credit estimates tend to produce noisier policy updates. Through analyses of critic quality, gradient noise, and update behavior, we investigate how Pass@ reweighting stabilizes single-rollout learning and why residual critic errors may limit SR-PPO’s ability to outperform GRPO.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.