When Agreement Labels Are Necessary for Worst-Group Model Selection
Abstract
Labels on points where all candidate classifiers agree cancel from average-risk comparisons, yet can determine the minimizer of worst-group error. We characterize when this agreement information is necessary for selecting among fixed binary classifiers. On a finite audit pool, a closed-form certificate gives the exact largest regret over a product of label-count intervals and supports adaptive querying and stopping through finite-population confidence sequences. Once all disagreement labels are known, an criterion determines whether a candidate meets the regret tolerance under every completion of the agreement labels. An indistinguishability bound establishes when no uniformly reliable agreement-free selector exists. Matching oracle bounds show that a second group can eliminate the quadratic label savings of small disagreement mass. CAWS uses the certificate with disagreement-first sampling and agreement fallback. Fixed synthetic, digit, Adult, and CelebA pools illustrate both sides of the criterion: four of ten CelebA banks exhibit an agreement obstruction. Matched-path ablations isolate the reduction in certification cost obtained by preserving shared losses.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.