Verifier-Induced Strategy Collapse: Family-Correlated False Negatives Eliminate Valid Strategies While Capability Metrics Improve
Abstract
When a terminal verifier in agentic reinforcement learning falsely rejects one strategy family, group-relative policy optimization (GRPO) does not merely mis-rank that family: it deletes it from the policy — in our experiments within about 30 steps — while every standard capability metric improves. The cause is the correlation structure of the checker's false negatives (FN), not their rate: recent theory models verifier noise as independent across solutions, an abstraction under which the deletion cannot occur. For GRPO with a terminal checker that rejects a fraction of valid rollouts of strategy family , we derive — in the tabular family-mixture abstraction — the exact finite-group dynamics: the expected one-step drift of a family's probability mass is up to a positive factor, where is the family's verifier acceptance rate and the group-mean acceptance, so selection reduces to a single threshold equal to mean verifier quality. The aggregate FN rate is provably insufficient to preserve strategy ordering, and a family goes extinct once . In a controlled tool-chain environment with an exact success oracle, a checker with family-correlated false negatives drove the targeted valid family from a third of the population (banking to ) to extinction within about 30 GRPO steps while true success rose from to and the checker's measured reliability converged to perfect. Matched-iid controls at the identical aggregate error rate, five-seed and 7B replications, a second model family, stress tests showing real capability loss on route-blocked and task-shifted environments, family-resolved audits of published GSM8K, HumanEval, and MATH-500 checker configurations support the theory and mark its scope.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.