Fooled at Probe Time, Ignored at Training Time: The Mechanism and Exploitation Boundary of Verifier Susceptibility in RL
Abstract
LLM judges used as reward signals for RL can accept wrong answers on some tasks—a property called verifier susceptibility. The natural response is to filter susceptible tasks from training. We show why that response fails, locate the gradient dose at which exploitation begins, and confirm both results with pre-registered experiments. The mechanism: in high-correlation regimes (where rollout-time proxy–gold correlation is positive, 45/45 training conditions), the policy gradient tracks the true objective. Susceptible tasks are structurally harmless—the fooling completions that define susceptibility at probe time are not the ones reinforced by optimization. The boundary: a pre-registered fixed-learning-rate dose ladder shows exploitation appears once roughly 26% or more of training groups carry reward signal (tested doses 6–65%)—below 7%, even a planted backdoor is ignored entirely—and doubling group size cannot buy signal, because token emissions are task-correlated. This locates verifier exploitation on the gradient-dose axis.The predictions, confirmed: four pre-registered replications show susceptibility-informed filtering is never detectably better than random task removal—and under a naturally low-correlation Mistral judge, a pre-registered equivalence test passes on the primary estimand. Susceptibility is also real (r=0.638 cross-prompt, at the reliability ceiling) yet no tested text classifier detects it beyond a difficulty baseline. We release the twenty-minute pre-flight probe that determines which regime a task–verifier pair is in.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.