Range-Constrained Disagreement Sampling for Label-Efficient Anytime Audits of Model Replacement
Abstract
Label-efficient model comparison does not automatically yield label-efficient safety certification. We study replacement of a frozen baseline by a member of a frozen classifier library on a fixed, initially unlabeled pool. For binary zero-one loss, the squared paired loss difference is exactly the observable prediction-disagreement indicator. This identity yields a label-free minimax second-moment design for all baseline contrasts. We characterize its simplex dual, its sharp gains on disjoint regions, and a no-gain regime in which one candidate covers the disagreement union. Because a variance-optimal proposal can have an unfavorable importance-weight range, we solve a probability-floor-constrained design by water filling and select its range using a label-free confidence envelope. We combine the proposal with established martingale tools, without-replacement residual estimation and deterministic completion bounds, obtaining simultaneous finite-pool certificates at every stopping time. Two experimental revisions contain 10,080 audit trials. In the final revision, the proposed design reduces mean maximum interval width by 5.8% against uniform-disagreement sampling on six real datasets at a 20% pool-label budget, but does not improve certification frequency and is worse in mean width than a strong predictable-mixture empirical-Bernstein comparator. The results identify a useful design characterization and a concrete variance-versus-certification failure mode, not universal empirical dominance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.