Audit Certainty Is a Signed Lever: Why a Verifier Calibrated on a Proxy Can Backfire in Deployment
Abstract
We calibrated a verifier on a proxy task, deployed it unchanged, and its strongest setting became its worst: raising audit certainty — the lever that cut violations on the proxy — raised them in deployment. Our claim — proxy calibration of audit certainty need not transfer across temporal structures for learning agents — rests on two findings of different standing. The first is derived. Classical deterrence says only the expected penalty, probability times severity, matters; for an agent that learns from the signal it does not, because rare severe penalties inflate the variance of the learned violation value. We prove an exact stationary-variance result and, under a Gaussian threshold approximation, a signed certainty law: raising certainty deters when the violation is value-suboptimal and backfires when it is value-optimal. The law predicted the inversion before the experiment; a deep-RL phase diagram confirms it, and the exact law bounds the approximation's error. The second finding is causally isolated but not derived, and the headline rests on it: the location of the sign boundary moves with task timing. With timing equalised it sits where the law predicts; delaying honest reward alone pushes it past every value gap we can measure, so a timing-mismatched proxy certifies the wrong policy. That displacement is orders of magnitude beyond the bounded approximation error, so the theory we prove does not explain it; we isolate the channel, rule out four candidate derivations and one whole route, and leave the closed form open. Under GRPO on an open LLM, over a pre-registered orthogonal grid of several hundred training runs, the certainty-over-severity ordering reproduces, and the logged advantages show why: group normalisation caps the advantage any single penalty can carry, discarding severity but not certainty, and a group-size sweep orders severity's weight as the bound predicts. Design rule: calibrate on matched timing, or measure the value gap under deployment timing before trusting a proxy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.