acceptodds
Under review as a conference paper at ICLR 2027

Identification Is Not Enough: Which Majority Examples Are Flagged Changes Last-Layer Correction

Abstract

Stratified last-layer correction fixes worst-group failures by upweighting examples a selector scores as likely minority, and selectors are compared by how well those flags recover the annotated minority. We test whether these summaries predict the outcome of the correction. Our matched control permutes, within each class, the scores of majority examples and leaves every minority example's score untouched. This provably fixes AUROC, precision, recall, the detected minority set and the expected weight each group receives, and changes only which majority examples carry the flagged weight. It needs group labels: a diagnostic, not a method. On CelebA, with two selectors (prediction confidence and a patch-blur score), random reassignment raises test worst-group accuracy on all five backbones under both balanced-sampling and expected-weight fitting. On the three original backbones the selectors' own assignments rank in the bottom 0–2.5% of 200 random reassignments on test, though higher (up to 44%) on two later backbones: the majority examples a selector flags, boundary cases frozen ERM often misclassifies, are worse to upweight than random ones. This is not only a threshold shift: after re-tuning each head's threshold on validation, reassignment still gains 2.7–3.2 worst-group accuracy points with no group losing more than 0.2, and a threshold-free worst-group-pair AUC agrees. Flagging the easiest majority examples helps about as much; fresh sampling seeds alone move accuracy under 1.4 points. The effect depends on the dataset: on Waterbirds and MetaShift, whose selectors flag few majority false positives, it is near zero under expected-weight fitting. At the default setting even a test-selected best reassignment trails frozen ERM; at validation-tuned settings random reassignment beats tuned class balancing in 7 of 10 backbone–selector pairs on test but only 5 of 10 on validation. Identification summaries and expected group weights are therefore not sufficient to predict the outcome of stratified last-layer correction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.