Predicting Reward Hacking in RLVR Requires Accounting for Gradient Geometry
Abstract
Reinforcement Learning from Verifiable Rewards (RLVR) has driven rapid LLM progress in math and coding, where verifying solutions is cheap. To move beyond these domains, RLVR is now commonly used with imperfect verifiers such as LLM judges. Yet training with imperfect verifiers can degrade final task performance as models learn to reward hack, i.e., to exploit verifier errors rather than fully solve the task. It is therefore desirable to predict the impact of verifier errors on task performance before committing substantial compute. Prior work suggests using confusion matrix statistics like Youden’s index. However, we prove that when verifier errors are systematic, these statistics are insufficient even for simple log-linear policies. Motivated by this limitation, we advocate for using information about the policy’s gradient geometry to improve predictions. We provide initial evidence about the promise of this direction through a simple gradient-based prediction procedure: we approximate the base policy’s tangent kernel using sampled generations and reuse this kernel to efficiently simulate linearized RLVR dynamics across many verifiers. All simulated policies can then be assessed on the same set of ground-truth labels, so benchmarking further verifiers is cheap. For models up to 8B parameters, we show that this linearized simulation already improves task performance prediction in many settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.