CIPS: Correct-Incorrect Prefix Score as a Baseline for RL with verifiable rewards
Abstract
Reinforcement learning with verifiable rewards (RLVR) is a simple and effective recipe for improving the reasoning capabilities of large language models, but its sequence-level rewards provide only a coarse signal for credit assignment. Cheaper alternatives to process reward models and Monte Carlo sampling use the likelihood of the correct answer as a prefix-level signal, but this likelihood also reflects spurious context and answer syntax rather than reasoning progress alone. We introduce the correct-incorrect prefix score (CIPS), which uses the likelihood of plausible incorrect answers from the same rollout group as a control for these effects. A prefix counts as progress only if it raises the likelihood of correct answers relative to incorrect ones. We show that CIPS is robust to spurious prefix content, retains the prompt-difficulty information of the group-mean baseline, and additionally distinguishes rollouts within a group and along the reasoning trace. Used as a prefix-dependent baseline in place of the group mean, CIPS requires neither a learned reward model nor additional rollouts. For two model scales and both mathematical and combinatorial reasoning tasks, CIPS consistently outperforms outcome-only training and prior dense-credit methods. Additionally, with essentially unchanged step time, it reaches the final performance of Dr. GRPO with about 25% fewer training steps. The code is publicly available (https://anonymous.4open.science/r/cips-65EB).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.