Certified Decision Transfer under Unlabeled Shift
Abstract
In content moderation and fraud detection, model upgrades should identify more positive cases without increasing review workload. Yet higher offline AUC does not guarantee this improvement under distribution shift, delayed target labels hinder evaluation before deployment, and changes in score scale complicate threshold reuse. We study certified decision transfer: using historical labels and current unlabeled traffic to decide whether a candidate should replace an existing policy. Rate Matching (RM) lets the reference determine how many cases to select and the candidate determine which cases fill those slots. We derive finite-batch and population regret decompositions that separate selection-count error from ranking error, and show why global AUC superiority is insufficient at a fixed review count. Certified RM (C-RM) uses paired policy disagreements on historical positives and bounds target selection-rate changes using unlabeled traffic. For pre-specified policies evaluated on independent certification samples, stable positive-case detection rates enable finite-sample control of erroneous approvals for each comparison, simultaneously for all ; insufficient evidence retains the reference. In a GoEmotions audit, 3,578 of 26,250 class-level upgrade comparisons are harmful; fully paired C-RM approves 834 with no observed population or finite-batch F1 degradation. On COCO, disagreement factorization certifies 4.0–4.6 times as many upgrades as the direct-cell test using the same data and nominal error budget. Our framework supports evidence-based model replacement before target labels become available. Code is available at https://anonymous.4open.science/r/CDT-86F8/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.