acceptodds
Under review as a conference paper at ICLR 2027

ProbePO: Sparse Probes, Redistributed Credit for Critic-Free LLM Reinforcement Learning

Abstract

Reinforcement learning from verifiable rewards assigns a single scalar reward to an entire reasoning trajectory, even though different parts of the trajectory can contribute very differently to success. GRPO converts that reward into one group-normalized advantage shared by every response token, while finer-grained approaches often rely on dense auxiliary rollouts or branching schemes that can also alter the data used for policy updates. We cast the missing middle as budgeted measurement: with a limited rollout budget, which partial solutions should be evaluated, and how should their noisy value estimates refine credit? ProbePO uses a reward-free uncertainty score to select a few partial solutions and estimates their expected terminal verifier reward from a capped set of continuation rollouts. ProbePO redistributes, rather than replaces, GRPO's trajectory-level advantage, changing where credit falls while preserving each trajectory's mean token-level advantage; the sampled continuations inform credit without entering the policy update. Across DeepSeekMath and Qwen-family models spanning 1.5B, 4B, and 7B parameters, ProbePO shows consistent observed improvements over GRPO: it is higher in all 18 benchmark–setting cells and is the highest evaluated method in 16 of 18. On Qwen3-4B-Base, ProbePO is 2.84 Avg-6 points above GRPO and 2.07 points above TreePO. In an end-to-end DeepSeekMath-7B comparison, ProbePO produced higher observed scores across all six benchmarks while using 76% fewer H200-hours than the released VinePPO recipe.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.