When Efficient Training Becomes Unfair: Auditing and Repairing Data Pruning
Abstract
Data pruning reduces training cost by removing examples a model appears to have learned, often with little change in average performance. However, average metrics can conceal substantial subgroup harm. First, we conduct a systematic demographic-fairness audit of data pruning across static and dynamic criteria and find that pruning can significantly degrade subgroup and worst-group performance even when removal is demographically uniform. This harm is not explained by preferential deletion of minority data; instead, pruning reduces overall training exposure, disproportionately affecting data-scarce groups. Motivated by this finding, we propose EquiPrune, a lightweight, modular wrapper that operates at the pruning decision stage to preserve sufficient training exposure. Across diverse datasets, architectures, and pruning strategies, EquiPrune restores baseline-level fairness in most affected settings while retaining substantial computational savings. These results suggest that efficient training should be evaluated on the fairness–efficiency frontier, not just accuracy–efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.