acceptodds
Under review as a conference paper at ICLR 2027

When Hard-Example Selection Fails: Gaussian Thresholds and Empirical Diagnostics

Abstract

Selecting the highest-loss examples can fall tens of accuracy points below random sampling at small retained fractions. We explain when this happens in a Gaussian model and test its predictions on image classification. For a fixed aligned score, we identify a threshold p*(r), given by one scalar equation in the signal-to-noise ratio r, and prove that with high probability in finite samples the selected class means invert below it, so that a nearest-mean classifier converges to the sign-flipped Bayes rule. The same population threshold holds for ridge-logistic regression. For a misaligned score we derive the selected means and the exact clean risk, so the sign flip becomes a continuous degradation with a computable above-chance boundary, and closed-form extensions cover independent score noise and a uniform-first random quota. On CIFAR-10, loss-ranked selectors do not invert real class geometry, but they displace class means four to six times farther than random selection does, consistently toward the nearest competing class. Two interventions on EL2N at a 5% fraction recover between a third and two thirds of its lost accuracy: scoring after 20 probe epochs instead of 1 raises final accuracy from 14% to 32%, and, in a separate paired experiment, replacing half of the hard quota with random draws under matched initialization raises it from 15% to 45%, against 62% for random selection. Only the second intervention moves the fixed-feature margin in step with accuracy. A geometry diagnostic that explains one successful intervention need not explain another, so such diagnostics should be validated across interventions before they are used to explain selection quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.