acceptodds
Under review as a conference paper at ICLR 2027

Measuring Damage, Not Learning: Step-0 Baselines for Continued Pretraining

Abstract

Data-recipe comparisons in continued pretraining share one design: continue from one checkpoint under competing rules, rank by held-out loss. Our case is latent-augmented pretraining, which inserts model-generated reasoning into training text uniformly: could a per-chunk signal place it better? We ran that comparison carefully — twelve conditions, up to three seeds, two domains, two model scales, a frozen truncation-safe metric — and obtained a clean, seed-stable ordering: uniform beat an equal-budget random control by nats/token, and demand-based routing lost to that control by . Then we scored the warmstart itself. Every condition, uniform included, is worse than the checkpoint it started from, so the ordering ranks damage, not learning. A tenfold lower learning rate removes of the headline contrast; with a warmstart that privileges no condition by density, both signs survive at – of their magnitude. Only replay plus a low rate puts runs back below their start, and there uniform loses its lead: random routing is ahead of it, significantly only when uniform alone matches the warmstart density, and no signal beats random. Across three public replications that share no code with ours ( runs), the learning-rate and two-axis step-0 checks reproduce every time; weight displacement, which tracked our ranking at , does not generalise. The mechanism, distribution shift and learning-rate annealing, is known (Wang et al., 2025); what is new is a documented case in which a ranking stable across seeds, domains and scales was a property of the recipe rather than of the allocations. Separately, length-corrected reasoning demand, a posterior-to-prior KL from a fixed model, predicts per-chunk latent benefit (; incumbent signal ). Of twenty-five published continued-pretraining data-recipe comparisons, most report the start, but loss-ranked ones more often as a figure point than as a number to difference against. Step 0 on two axes and a rate check cost an evaluation each; we argue they belong in the standard protocol.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.