Feasible-Source Rules Define Dataset-Mixture Evaluation Identity: Cross-Schema Fixed-Pool Interventions
Abstract
Matching a raw source pool leaves a critical attribution variable unresolved in dataset-mixture studies: the feasible-source rule that determines which sources enter allocation. Changing this rule can redirect training mass and alter benchmark conclusions even when the model, token budget, allocator, and evaluation harness remain fixed. We formalize rule-conditioned evaluation identity and introduce a counterfactual support-intervention audit that isolates rule-induced displacement along the rule-to-mask-to-mixture-to-score path. In 100B-token Llama-3 8B continued pretraining, SetTrace reaches 66.8% under the research rule and 63.2% under redistributed weights, with 2.1% and 1.4% proxy violation, respectively. Across three predeclared redistributed-weights encodings, 16–35% of mixture mass moves and Holm-corrected winner decisions change on HumanEval, GSM8K, and TriviaQA. Independent reconstructions with newly authored admissibility semantics, separately tuned SetTrace, constrained DoReMi, and masked DoGE instruments, and a separately annotated 96-unit pool place research-to-redistribute displacement at 2.8–3.6 pp, with every paired 95% interval excluding zero. A blind six-counsel re-annotation preserves the SetTrace-over-Manual-Curator ordering by 3.8 pp and the same winner-instability pattern, while a disjoint FineWeb/Dolma/StarCoder2–Mistral-7B replication retains a 4.1 pp gap. These results make feasible-source matching a condition for identifying allocator effects and establish rule reporting as part of the dataset-mixture evaluation specification.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.