Better Policies, Worse Verifier Audits
Abstract
Verifier stacks increasingly sit inside post-training and test-time compute loops. We study a specific control loop: a component verifier is evaluated on fresh policy outputs against the full success label and refreshed when that aggregate score falls. For a conjunctive target , its on-policy AUC is , where . As the policy masters , can rise toward one, removing weight from -failures and rewarding information about . We call this on-policy audit collapse. We derive a sharp identified set, a complete-erasure threshold, and the exact point at which the audit can prefer a worse component verifier; its population optimum is the full-outcome predictor, not the component oracle. An exact oracle exhibits the collapse, a natural multi-label text control shows inversion can select a head with lower canonical -AUC at the observed failure mixture, and a prediction-registered ECG study shows that acting on the audit can reduce off-policy -AUC, transfer the damage to an unseen generator, and increase downstream constraint violations. Preserving own-failure support prevents the damage. Matching label-complexity bounds and a controlled-mixture audit give the corresponding repair: audit the component on its own failures, and keep those failures in the refresh distribution.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.