acceptodds
Under review as a conference paper at ICLR 2027

Are Coreset Selection Methods Worth Their Cost?

Abstract

Data selection methods, often called coreset selection, pick a training subset to make training cheaper. The usual evaluation reports downstream accuracy at a fixed subset size, which leaves two quantities out of the metric: the time spent selecting the subset, and the training recipe behind each reported number. We close that gap with an end-to-end benchmark that standardizes downstream training and prices selection and training on one wall-clock axis estimated from controlled timing probes, spanning 4 datasets from CIFAR-10 to ImageNet-1K, 11 selectors, 5 fractions, and 3 seeds, across 1,674 released runs with their selected indices. Repeated-sampling work has shown that budget-aware evaluation already favors random strategies, and our two budget studies test whether that holds when every selector is granted its most favorable tested operating point. Across eight budget anchors on each of CIFAR-10 and Tiny ImageNet, no anchor is won by a sophisticated selector: every winner is class-balanced random sampling, repeated random sampling, or full-data training. In nearby-cost comparisons on ImageNet-1K, training on all data for fewer epochs has the highest accuracy among the operating points we probe. A per-dataset cost audit shows that selection cost is dominated at every scale by a fixed full-dataset scan, so it cannot be amortized away by selecting a smaller fraction, and its absolute size does not extrapolate from one dataset to another. We further separate cost amortization at a fixed architecture from cross-architecture accuracy transfer, and document 9 correctness fixes to a widely used codebase, one of which shifts a standard Herding baseline by 5.9 percentage points (pp). Selection time is not free preprocessing, and an evaluation that ignores it measures the wrong quantity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.