What Does 5% Mean? The Coreset Regime Boundary Is Pinned in Samples per Class
Abstract
Coreset selection methods are usually evaluated at fixed fractions of the training set, yet the winning selection principle changes with budget. Where does this change occur, and which budget coordinate makes it stable? We measure the crossover budget m0 between a learnability probe (LFrac) and a coverage anchor (Herding), observing the reversal directly on four datasets including ImageNet-1k. Two controlled interventions distinguish the effects of pool size and visual bandwidth under fixed selection and training protocols. Shrinking the per-class pool of ImageNet-100 by 8.4× leaves m0 within 65–80 samples per class while its fraction reading drifts by roughly 10×; CIFAR-100 reproduces this approximate stability. Reducing effective pixel count by 49× on a fixed input canvas lowers full-data accuracy by 7.4 points, while m0 shifts systematically from about 66 to 80 within the same sampled budget bracket. The strongest invariance is therefore to pool size: the crossover behaves like an approximately fixed per-class sample requirement of the measured selection and learning system. This is more than unit conversion, because matching fractions can place methods in different budget regimes. Locating the boundary without downstream measurement remains open: the six dataset descriptors examined here provide no validated predictor, and naive subset mixtures on CIFAR-100 do not reliably beat the better pure endpoint. Oracle-tuned CCS improves on both endpoints near the transition, but a fixed setting transferred to ImageNet-100 loses at the smallest budget. These findings motivate reporting absolute samples per class and pool size alongside fraction in cross-dataset evaluations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.