If You Can't Beat It, Copy It: Hint Reliance and Learning Signals in RLVR
Abstract
Reinforcement learning with verifiable rewards (RLVR) checks only the final answer. When a prompt contains a candidate solution, a policy can earn the same reward by solving the problem or by copying the candidate. We study which route RLVR strengthens, using Countdown problems with hints of controlled reliability. Checkpoints trained with always-correct hints reach near-perfect correct-hint accuracy yet range from 0.00 to 0.60 in wrong-hint accuracy. In Countdown with Qwen2.5-7B and Llama-3.1-8B-Instruct and in an arithmetic task with Qwen2.5-Math-7B, copying grows when the hint is more reliable than the accuracy the model reaches without hints at the same budget. A model that cannot beat the hint learns to copy it, and the updates spent copying are lost to learning. Copying persists because a saturated copier gives every response to a prompt the same reward, which leaves group-relative RL no contrast within a prompt. Adding no-hint questions restores this contrast. The first response token can choose between the routes a model has. On held-out questions, replacing one hidden state lowers wrong-hint copying from 0.97 to 0.35 and from 0.80 to 0.28, and forcing only the first token lowers it to 0.32 and 0.31. The same control applied during training creates reward contrast only while it is on. RLVR learns from the alternatives that the policy itself samples.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.