Pick the Mixture, Not the Samples: Combinatorial Data Selection for LLMs
Abstract
Data selection for LLM fine-tuning faces two facts about real training: data pools draw from many distinct domains, and evaluation spans diverse tasks. As a result, the same sample helps different tasks differently, and samples interact during training so that a subset’s value is not the sum of its parts. Current data-selection methods leave these realities largely unaddressed. We formalize them as Combinatorial Data Selection (CDS), a framework that separates attribution (how samples are scored) from selection (which subset to pick) and models the non-additive interactions among samples. On a real instruction-tuning data pool, we find substantial cross-domain complementarity (joint utility exceeds the sum of singletons), alongside redundancy within domains. Inspired by this, we propose Mixture-then-Select, a two-stage algorithm that first decides how many samples to take from each domain, then ranks samples within each domain by attribution score. Theoretically, under the stated structural and predictor-accuracy assumptions, the gap to the optimal subset is linear in the selection size and independent of data pool size. Empirically, Mixture-then-Select outperforms strong baselines by 20 to 54% in held-out utility, is robust across attribution methods, and works across different models and data pool sizes. The framework also extends to the data purchasing problem under a budget constraint when sample prices vary. Our findings suggest that allocating selection across domains offers a promising path toward more efficient and effective data selection for modern LLMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.