Tied, Not Won: Testing Claims That One Model Scales Better Than Another
Abstract
Many papers conclude that one architecture, training objective, or pretraining corpus "scales better" than another. Because a scaling rate is the exponent of a fitted power law, such a conclusion is a claim about the difference between two exponents. In practice it is drawn from a handful of model sizes, and neither the exponents nor their difference come with a standard error. We supply the missing statistics. Using only the per-scale results that the original papers released, we refit each power law, check that the refit reproduces the published exponent, and test whether the gap between two exponents is distinguishable from zero. Across widely cited comparisons from eight released model suites, most claimed differences in scaling rate are statistical ties. The conclusion that the vanilla Transformer scales best among ten architectures does not hold against its close contenders, and nearly all pairs of pretraining corpora in the DataDecide suite differ in loss level but not in rate. The probability that the claimed winner would reverse, computed from the same test, separates these ties from large, well-established differences. A closed-form bound on the smallest detectable gap explains the pattern: the small sets of model sizes in common use cannot resolve gaps of the size being claimed. A survey of recent scaling papers further shows that most do not release the data needed to check their own comparisons. We close with simple reporting practices that make comparative scaling claims testable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.