acceptodds
Under review as a conference paper at ICLR 2027

Let Intuition Teach Deliberation: Rewarding Prefix Success Potential in Reasoning Models

Abstract

Two reasoning trajectories can reach the same answer while differing in whether a correct short answer was already available halfway through. We use this distinction to let verified short-answer probes supervise long-form deliberation. Prefix Success Potential (PSP) measures the probability of producing a correct short answer from a prefix under a fixed protocol. The same policy samples auxiliary answers, a verifier scores them, and their success rates reward the complete original trajectory. Our variant PSP-ME (ME: midpoint and endpoint) favors higher midpoint potential at equal endpoint potential. Probes serve only for scoring; no external teacher or process annotations are required. We train with an 8,192-token budget and evaluate at 32,768/38,912 tokens. Across Qwen3-1.7B, Qwen3-8B, Gemma4-E4B-it, and Gemma4-12B-it, PSP-ME achieves Math-7 scores of 51.06%, 68.70%, 65.56%, and 79.93% respectively. These represent gains of 3.89, 4.45, 5.55, and 6.56 percentage points over matched outcome-reward GRPO. Fixed-prefix evaluation reveals two key findings. First, RL training dissociates answer potential from natural completion: Gemma4-12B-it improves potential by 14.32 pp while completion drops by approximately 29 pp, suggesting that PSP teaches models to maintain extractable reasoning states rather than forcing earlier generation. Second, gains transfer heterogeneously through different mechanisms: Gemma models improve short-answer potential while Qwen3-8B does not, yet both achieve substantial long-budget gains, demonstrating that midpoint supervision enhances reasoning quality through multiple pathways rather than a single uniform mechanism.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.