acceptodds
Under review as a conference paper at ICLR 2027

Measuring Uneven Class Damage from Data Pruning: Four Pitfalls and a Control Ladder

Abstract

Score-based data pruning damages some classes more than others, and the standard way of measuring that damage cannot tell three explanations apart: the class imbalance a difficulty score induces, the choice of *which* examples survive inside each class, and the plain fact that a pruned model is less accurate. We document four measurement pitfalls that produce false positives here, each demonstrated on a conclusion of our own that we then withdrew: a disagreement statistic reported without its mechanical floor, which inflated a churn ratio to 2.08 when the floor accounts for essentially all of it; a baseline fitted through the very group under test; two non-pruning control families that disagree at matched accuracy; and selection-seed variance as large as training-seed variance, in a setting where 16.3% of a canonical Forgetting coreset is decided by score ties and only 94.3% of its indices are stable across seeds. Three of the four apply to any study that measures how much two models disagree. We then build a ladder of five increasingly strict controls, fixing budget, then per-class counts (Gini), then the hard/random mixing fraction q, then final accuracy, and run the same question up it. An aggregate hypothesis dies at the second rung: against a non-pruning control trained to the same accuracy, the difference is +0.0005, CI [-0.023,+0.023]. A class-level hypothesis survives all five on the setting where we found it, bracketed at [-1.26,-0.71] by two comparisons that fail in opposite directions, and persists under a second architecture. It does not generalize: borderline on Food-101, null under real human label noise, and sign-reversed on Tiny-ImageNet with an interval that excludes zero. Subsampling CIFAR-100's test set to Tiny-ImageNet's per-class count leaves the effect intact in 99% of draws, so the reversal is a boundary of the mechanism and not small-sample noise. We report the ladder, the failures, and the pitfalls as the portable contribution.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.