acceptodds
Under review as a conference paper at ICLR 2027

Green Because They Leaked: Test-Gated Policy Updates Don't Certify Safety

Abstract

Automated judges—compiled rule systems, fine-tuned classifiers, or prompted language models—increasingly gate what AI agents are allowed to do. When the policy changes, the judge must be updated, and the standard safeguard is borrowed from continuous integration: admit an update only if it passes a suite of test cases. We show that this green checkmark means far less than it appears to. First, the standard way of measuring the safeguard leaks. When gating tests and evaluation cases share underlying evidence states, the gate is credited with catching errors on cases it has already screened: under such a protocol, our test-gated updater appears to buy monotone safety, reaching 100% error catch at large test budgets; under a corrected state-disjoint protocol, error catch saturates at 21% from budget one, and we explicitly retract the earlier result. Second, with leakage removed, safety is bounded by coverage, not budget. The effective budget is the number of unique evidence states that exist—about three per revision on our held-out split, zero for a third of revisions—and every admitted erroneous update lies on a criterion with no test coverage; neither test quality, gate strictness, nor prompted model scale repairs the gap. We then address both problems: the leakage, by the state-disjoint protocol itself, released with validators; the coverage, by a proven converse to our bound—checking the candidate against the archived old rule on the full typed state space outside the amendment's scope certifies zero protected-side errors, with no labels required—realizable by schema-based test generation, whose empirical validation is left open. Until then, inventory-drawn test suites make update gating an auditable screen, not a safety certificate. We release the benchmark, partitions, and audits.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.