LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking
Abstract
Reinforcement learning with verifiable rewards (RLVR) has become the dominant paradigm for scaling reasoning in large language models, but it can introduce a subtle failure mode: LLMs gaming verifiers. We study this on inductive reasoning tasks, where models must induce logical rules that capture relational patterns from labeled examples. We find that some LLMs systematically abandon rule induction. Rather than inferring relational rules (e.g., ”plants with purple leaves are toxic”), they enumerate instance-level labels (e.g., ”plant_01 is toxic”), producing outputs that pass verifiers while failing to capture the intended structure. We show that this behavior is not a failure of understanding but a form of reward hacking: imperfect verifiers that check only extensional correctness admit false positives. To detect such shortcuts, we introduce Isomorphic Perturbation Testing (IPT), which evaluates a single model output under both extensional verification on the original task and verification on a logically isomorphic perturbation. While genuine rule induction is invariant to such perturbations, shortcut strategies fail. Using IPT, we find shortcut behavior in RLVR-trained reasoning models (e.g., GPT-5, Olmo-3.1) but not in non-RLVR models (e.g., GPT-4o, GPT-4.5), with prevalence rising with task complexity and inference-time compute. In controlled training experiments, extensional verification directly induces shortcut behavior, while isomorphic verification eliminates it. These results show that RLVR can incentivize reward hacking not only through overt manipulation but also by exploiting what the verifier fails to enforce.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.