What Do Imperfect Verifiers Teach Models? Persistent Susceptibility to Reward Hacking
Abstract
Verifiers supply rewards, select answers, and filter the data used to train new models. Scaling reliable and robust verification is a bottleneck for real-world uses of LLMs, yet practical verifiers make systematic errors. When a verifier is biased, what does it teach a model, and what does that model pass on? We study RL on arithmetic answered in text, with a verifier that also rewards wrong answers containing a marker token (e.g., python), and find that it leaves behind a susceptibility: the probability that further training with the verifier ends in reward hacking, as exact-reward accuracy collapses. In controlled experiments, (1) susceptibility is invisible in current metrics: models with the same accuracy differ in whether further RL makes them collapse. (2) It persists: after fine-tuning on correct solutions, a model that once exploited the verifier collapses again, unlike one trained under a verifier keyed to another token. (3) It barely spreads: teacher outputs filtered of the marker pass on the habit of writing it, but little of the susceptibility. (4) Credit controls it: restoring RL credit to correct answers prevents collapse, even after such a history. These findings advance the understanding of RLVR and expose a challenge for alignment: vulnerabilities a verifier leaves behind escape tests of current performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.