From Selection to Exposure: Online Data Selection via Learning-Signal Propagation
Abstract
Data selection aims to reduce the cost of training by allocating computation toward informative examples, while maintaining comparable performance as example utility changes throughout learning. Existing methods estimate sample importance using early-training statistics, losses, or learned selection policies and then prune or prioritize a subset of the training data. While effective, these approaches can over-concentrate training on currently prioritized samples, reducing coverage of other regions of the training distribution. We propose an online data-selection framework that turns each mini-batch into feedback for selecting subsequent training examples. At each optimization step, feedback from the current batch continuously reshapes the sampling distribution over nearby examples in the evolving representation space, increasing exposure to currently informative regions without deterministically excluding lower-priority samples. This allows the current batch to directly influence future training exposure while preserving broader data coverage and reusing representations naturally produced during training. Across image classification, semantic segmentation and instruction tuning, SPADE achieves strong predictive performance with high training efficiency. On Tiny-ImageNet, SPADE reaches comparable accuracy with up to 67.2% fewer training sample presentations and 64.1% less wall-clock time. Across the three image-classification benchmarks, it achieves the best Top-1 accuracy in 12 of the 15 settings evaluated. On ADE20K, SPADE reaches 42.0% mIoU using 70% of the training data, improving over the compared vanilla baseline. For LLaMA-7B instruction tuning using only a small fraction of Alpaca, SPADE also improves over the compared baselines, achieving the highest average score of 26.5 across BBH, DROP, MMLU and HumanEval.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.