Works Where Calibrated, Fails Where Misused: Auditing Small-Scale Pretraining's Seed-Noise Instruments
Abstract
Small-scale pretraining selects data recipes under a handful of seeds, read by three instruments. On public per-seed assets (plus two small self-run controls; Appendix H), we audit each where it is used: two of the three reproduce at home (the floor's per-seed data were never released), and each fails where misused — differently for each. The checkpoint-noise proxy of Signal & Noise reproduces at home, but its advertised correlation is beyond certification at realistic seed budgets under the conventional Fisher- certificate (a perfect proxy reads – against ), and it under-reads seed noise in magnitude on the accuracy readout ( median, up to , reversed on perplexity). An init-only noise floor cannot identify the all-source floor from single-source arms at any budget; only a bundled or crossed initorder design can. DataDecide's recipe ranking reproduces its exactly at its recommended scale, but on the accuracy readout its deconvolved separability SNR stays at or below up to 60M, and at 4M the macro ranking inverts against the 1B ranking (; robust to recipe resampling, borderline once seed noise is priced in) and seed averaging does not rescue it; the 11-domain perplexity readout does not invert. We release code, a per-cell noise atlas, and report label–config-mismatched files in PolyPythias's release.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.