acceptodds
Under review as a conference paper at ICLR 2027

If You Can't Beat It, Copy It: Hint Reliance and Learning Signals in RLVR

Abstract

Reinforcement learning with verifiable rewards (RLVR) checks only the final answer. When a prompt contains a candidate solution, a policy can earn the same reward by solving the problem or by copying the candidate. We study which route RLVR strengthens, using Countdown problems with hints of controlled reliability. Checkpoints trained with always-correct hints reach near-perfect correct-hint accuracy yet range from 0.00 to 0.60 in wrong-hint accuracy. In Countdown with Qwen2.5-7B and Llama-3.1-8B-Instruct and in an arithmetic task with Qwen2.5-Math-7B, copying grows when the hint is more reliable than the accuracy the model reaches without hints at the same budget. A model that cannot beat the hint learns to copy it, and the updates spent copying are lost to learning. Copying persists because a saturated copier gives every response to a prompt the same reward, which leaves group-relative RL no contrast within a prompt. Adding no-hint questions restores this contrast. The first response token can choose between the routes a model has. On held-out questions, replacing one hidden state lowers wrong-hint copying from 0.97 to 0.35 and from 0.80 to 0.28, and forcing only the first token lowers it to 0.32 and 0.31. The same control applied during training creates reward contrast only while it is on. RLVR learns from the alternatives that the policy itself samples.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.