acceptodds
Under review as a conference paper at ICLR 2027

CIPS: Correct-Incorrect Prefix Score as a Baseline for RL with verifiable rewards

Abstract

Reinforcement learning with verifiable rewards (RLVR) is a simple and effective recipe for improving the reasoning capabilities of large language models, but its sequence-level rewards provide only a coarse signal for credit assignment. Cheaper alternatives to process reward models and Monte Carlo sampling use the likelihood of the correct answer as a prefix-level signal, but this likelihood also reflects spurious context and answer syntax rather than reasoning progress alone. We introduce the correct-incorrect prefix score (CIPS), which uses the likelihood of plausible incorrect answers from the same rollout group as a control for these effects. A prefix counts as progress only if it raises the likelihood of correct answers relative to incorrect ones. We show that CIPS is robust to spurious prefix content, retains the prompt-difficulty information of the group-mean baseline, and additionally distinguishes rollouts within a group and along the reasoning trace. Used as a prefix-dependent baseline in place of the group mean, CIPS requires neither a learned reward model nor additional rollouts. For two model scales and both mathematical and combinatorial reasoning tasks, CIPS consistently outperforms outcome-only training and prior dense-credit methods. Additionally, with essentially unchanged step time, it reaches the final performance of Dr. GRPO with about 25% fewer training steps. The code is publicly available (https://anonymous.4open.science/r/cips-65EB).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.