acceptodds
Under review as a conference paper at ICLR 2027

Feasible-Source Rules Define Dataset-Mixture Evaluation Identity: Cross-Schema Fixed-Pool Interventions

Abstract

Matching a raw source pool leaves a critical attribution variable unresolved in dataset-mixture studies: the feasible-source rule that determines which sources enter allocation. Changing this rule can redirect training mass and alter benchmark conclusions even when the model, token budget, allocator, and evaluation harness remain fixed. We formalize rule-conditioned evaluation identity and introduce a counterfactual support-intervention audit that isolates rule-induced displacement along the rule-to-mask-to-mixture-to-score path. In 100B-token Llama-3 8B continued pretraining, SetTrace reaches 66.8% under the research rule and 63.2% under redistributed weights, with 2.1% and 1.4% proxy violation, respectively. Across three predeclared redistributed-weights encodings, 16–35% of mixture mass moves and Holm-corrected winner decisions change on HumanEval, GSM8K, and TriviaQA. Independent reconstructions with newly authored admissibility semantics, separately tuned SetTrace, constrained DoReMi, and masked DoGE instruments, and a separately annotated 96-unit pool place research-to-redistribute displacement at 2.8–3.6 pp, with every paired 95% interval excluding zero. A blind six-counsel re-annotation preserves the SetTrace-over-Manual-Curator ordering by 3.8 pp and the same winner-instability pattern, while a disjoint FineWeb/Dolma/StarCoder2–Mistral-7B replication retains a 4.1 pp gap. These results make feasible-source matching a condition for identifying allocator effects and establish rule reporting as part of the dataset-mixture evaluation specification.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.