Towards Theory-Guided and Generalizable Data Selection
Abstract
Data selection aims to construct compact coresets that reduce training cost while preserving model performance. Most existing methods select data by preserving statistical, geometric, or optimization properties of the full training set, reflecting an empirical risk minimization (ERM)-oriented perspective. However, discrepancies between observed training data and unseen test data can arise from dataset collection, construction, and sampling effects, suggesting that preserving full-data ERM behavior alone may be insufficient for generalization. We therefore revisit data selection from a generalization perspective. Drawing on a domain generalization error bound, we identify four key factors associated with generalization and formulate corresponding principles for data selection: preserving predictive utility, encouraging cross-subset invariance, maintaining semantic consistency, and promoting distributional coverage. Guided by these principles, we propose GenDS, a theory-guided Generalizable Data Selection framework that operationalizes them through three complementary sample-level scores and a coverage-aware coreset selection strategy. Specifically, the scoring stage evaluates image-text alignment, subset invariance, and classification-structure consistency to capture predictive utility, cross-subset invariance, and semantic consistency, respectively, while the subsequent selection stage promotes distributional coverage and reduces redundancy. Extensive experiments on six standard and multi-domain image classification benchmarks demonstrate that GenDS consistently outperforms representative data selection methods across varying selection ratios, with particularly strong gains under limited data budgets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.