LACS: Distribution-Sensitive Label-Agnostic Confidence Selection for Dynamic Data Pruning
Abstract
Training on large datasets spends much of its computation on examples whose value changes as the model learns: confidently learned examples contribute little further gradient signal, while confident mistakes can dominate the next fitting step. Dynamic data pruning reduces this cost, but a fixed quota decides how much to prune before seeing the current score distribution, and an absolute threshold is brittle to the scale and drift of model confidence. We introduce LACS (Label-Agnostic Confidence Selector), a training-time subset selector whose threshold and retained count are both determined by the current confidence distribution. LACS reads only each example's maximum predictive confidence—never its label—and removes an example when its confidence exceeds a log-mean-exp reference of the current population plus a margin. At a fixed mean confidence this reference moves with dispersion and upper-tail shape, and it admits an exact no-pruning criterion: nothing is removed when too much confidence mass lies near the current maximum. The pruned fraction never exceeds a user-set bound. On real training batches the rule behaves as the analysis predicts: the pruning rate rises from about 11% early in CIFAR-10 training to 43% in mid-training and returns to 2% once confidences become uniformly high, the no-pruning condition matches the observed decision on every logged batch, and batches with equal mean confidence but different shape receive different retained fractions. We evaluate LACS in seven settings spanning neural, language-model, long-tailed, and semi-supervised training and three repeated-fitting (boosting) regimes. In neural training LACS removes 10–35% of examples from backward propagation with accuracy within seed noise of full training on CIFAR-10 and CIFAR-100; at a matched backward budget and matched number of optimizer updates it outperforms small-batch training by 1.0 pp and random selection by 1.2 pp over five seeds, and it has the higher mean than every work-matched dynamic-pruning baseline we ran. In repeated fitting the selected support improves prediction: across 14 tabular datasets, LACS lowers the mean test error of a quadratic boosting learner from 23.36% to 22.96% while fitting about 30% fewer rows over full replays, and it beats a count-matched random support on 13 of 14 datasets. In fitting-dominated RBF-SVM boosting the saved work becomes wall-clock savings of 30–94% in seven of eight dataset/split cells, and none where the rule prunes nothing. Executed work and wall-clock time are reported separately: the neural implementation adds a scoring forward and does not yet convert its sparsity into faster training, and cached scores alone do not change that; the conversion is a systems question, not a property of the rule.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.