Selection Is Not Validation: Auditing Proxy-to-Behavior Reversals in Language-Model Data Curation
Abstract
Data-selection methods for language-model training are often validated with the same proxy used to construct the training set. This creates a circular evaluation shortcut: a selector can improve its target score without improving the trained model. We introduce a staged claim-advancement audit that treats selection scores as hypotheses rather than outcomes. It specifies random and credible simple-baseline comparisons, audits training parity, and organizes source-clustered uncertainty, paired reference-loss and generation analyses, retention checks, and mechanism diagnostics, while marking missing evidence explicitly. Unlike a pass/fail checklist, the audit maps each result to the strongest claim supported by its evidence and identifies the next experiment required to advance that claim. In linked discovery cases in mathematical pretraining and instruction tuning, stored evaluation records reveal source-dependent uncertainty, aggregation-dependent likelihood estimates, and disagreements between reference likelihood and decoded reference overlap on shared example IDs. These paired analyses and a single-annotator, conflict-enriched 20-pair calibration provide bounded diagnostic evidence; useful fitted-model results are retained even when broader claims remain unresolved. The resulting evidence map turns proxy optimization into an actionable validation program: reject unsupported downstream claims, qualify surface-specific gains, or advance only those effects that survive independent behavioral and mechanism tests.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.