Measure, Don't Prescribe: Inferring Budget-Dependent Difficulty Preference for Dataset Pruning
Abstract
Dataset pruning selects a compact subset to reduce the cost of training on large datasets. Which examples are worth keeping depends on pruning budget: small budgets favor easier, representative examples, whereas larger budgets increasingly favor harder, more informative examples. Existing methods first prescribe a parametric form for the difficulty preference (e.g., hard cutoff, Beta distribution, or contiguous window), and adapt across budgets by adjusting its parameters (e.g., a cutoff ratio, shape parameters, or window position). However, identifying the optimal parameter values at the target budget requires trial-and-error search, and the form is fixed in advance, which may not suit every setting. Instead, we propose Difficulty Preference Scoring (DPS), which replaces prescribe-then-search with probe-then-infer: DPS first constructs a difficulty-balanced probing subset, then extracts from its training dynamics the signals that determine how the budget should be adjusted across difficulty levels, and corrects the allocation accordingly. This reduces the adaptation cost at a target budget from extensive trial and error to a single probing run, and lets the preference be inferred from the target setting rather than prescribed in advance. Across a wide range of budgets (1%–50%) and diverse datasets (CIFAR-10, CIFAR-100, QNLI, ImageNet and SQuAD 1.1), DPS achieves superior performance over state-of-the-art pruning methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.