acceptodds
Under review as a conference paper at ICLR 2027

Recycle Then Drop: Selective Recycling for Efficient Data Pruning

Abstract

Training larger models on more data has improved performance but also increased training cost. Dynamic data pruning reduces this cost by scoring each sample with the current model and excluding those the model already handles easily. These pruning scores are typically computed from augmented inputs. However, a score under one augmentation does not establish that the sample has little training value under another. Direct exclusion can therefore leave useful training signal unused. To address this, we propose Recycle Then Drop (RtD), which selects a fraction of the exclusion candidates and trains them with stronger augmentation. This extends the binary pruning decision to keeping, recycling, or dropping samples. Because these decisions still require scoring every sample, we further introduce an amortized variant, A-RtD, which periodically updates and reuses them to reduce repeated scoring and input processing. Experiments across five backbones, multiple datasets, and various augmentation settings demonstrate improvements in accuracy and training efficiency. On CIFAR, RtD reduces the number of samples used for backpropagation by over 80%, while A-RtD reduces training time by up to 54%, both achieving higher accuracy than full data training. Both methods also outperform full data training on object detection and semantic segmentation while requiring less backpropagation. Additional evaluations demonstrate compatibility with existing pruning methods. These results establish selective recycling as a practical approach to improving the accuracy and efficiency of data pruning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.