acceptodds
Under review as a conference paper at ICLR 2027

When Can Old Evaluations Certify a New Model? Matching Label-Complexity Bounds Under Evaluator-Residual Drift

Abstract

Releasing a model update requires certifying that its risk on the current population stays below a threshold. Trusted labels are expensive, whereas a cheap evaluator, such as a language-model judge, scores every example. Reusing evaluator errors from earlier audits is tempting, yet it is unclear when such evidence may replace current labels. The answer depends on the status given to history. If the errors can change invisibly, no label-free test detects the change, and every valid, useful certifier must keep buying labels at a rate we characterize; if a bound on the change is assumed, label-free certification is valid at an explicit error cost. For the middle ground, where history is informative but untrusted, we propose portfolio vigilance, a sequential certifier that mixes a betting expert guided by history with one that learns only from current labels; history affects only how it bets, so validity holds for any history. The contribution is not prior-informed betting or expert mixtures themselves, but the distinction between historical evidence that may enter validity and historical evidence that may only guide evidence accumulation. In a canonical model, accurate history shortens decisions but never raises the evidence growth rate, whereas stale history can destroy it. On held-out CIFAR-10N and DICES-990 data, portfolio vigilance needs (95% confidence interval (CI) ) and () times the labels of a matched prediction-powered monitor with no observed false certification, and fewer labels on all six external blocks. When historical advice is corrupted, it stays within of its better component, whereas trusting history alone costs up to times as much. In post-confirmatory repeated-judge experiments on DICES-990 and ToxicChat, changing the rubric of a fixed LLM judge moves its scores more than its run-to-run variation. After the change, the portfolio needs () and () times the labels of the matched monitor, and fewer labels than trusting history alone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.