Evolved Sampling: Efficient Training through Loss-Dynamics-Aware Data Selection
Abstract
Data selection aims to accelerate training while preserving model performance. To achieve this, a common approach is to identify informative data samples with significant contributions to the training. In this work, we propose **Evolved Sampling** (**ES**), a simple yet effective framework for *dynamic* sampling along the training process. This method conducts batch level data selection based on not only the dynamics of losses but also *correlations between gradients* along the training, while involving lightweight *re-weighting computations only regarding losses*, significantly reducing the back propagation time with maintained model performance. Due to its conciseness, ES is also readily extensible to incorporate set level data selection (to form ES with pruning, **ESWP**) for further accelerations. As a plug-and-play framework, ES(WP) is mathematically proved to converge, and consistently achieves lossless training accelerations across multiple pre-training and post-training tasks with various scales, saving up to 43.7% wall-clock time. Our results motivate further investigations on the data efficiency aspect of modern large-scale machine learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.