acceptodds
Under review as a conference paper at ICLR 2027

RLVR as Prompt Augmentation reinforces both Correct and Incorrect Reasoning to the Correct Answer

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR)-based post-training of Large Language Models (LLMs) has been shown to improve accuracy on reasoning tasks and continues to attract significant attention. While RLVR is known to improve solution accuracy, there has been little agreement on its effect on the so-called reasoning traces–with previous investigations coming to differing conclusions. We argue that outcome rewards cannot distinguish between reasoning from any other scaffold that leads to final answer correctness. To support our claims, we show that RLVR as learning prompt-augmentation conditioned on solution correctness, where a CoT is reinforced when it raises the probability that the model produces a correct final answer independent of whether CoT is correct or not. The objective is indifferent to every other property of the CoT including semantics, interpretability and faithfulness. We derive testable predictions and evaluate the prior results in a formally verifiable domain and test on GSM8K and two of its variants, GSM8K-Symbolic-P1 and GSM8K-Symbolic-P2. For GSM8K, where trace validity cannot be reliably established, we measure trace coherence, defined by the absence of errors. We compute , the probability of a correct final answer given a coherent trace, and , the probability given an incoherent one, for the base model and after RLVR training. Across all settings rises without exception, contradicting the claim that RLVR implicitly incentivizes correct reasoning. Thus RLVR learns prompt augmentations conditioned on final answer correctness: the objective does not specifically upweight valid reasoning, but any CoT that increases final-answer correctness.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.