acceptodds
Under review as a conference paper at ICLR 2027

Tied, Not Won: Testing Claims That One Model Scales Better Than Another

Abstract

Many papers conclude that one architecture, training objective, or pretraining corpus "scales better" than another. Because a scaling rate is the exponent of a fitted power law, such a conclusion is a claim about the difference between two exponents. In practice it is drawn from a handful of model sizes, and neither the exponents nor their difference come with a standard error. We supply the missing statistics. Using only the per-scale results that the original papers released, we refit each power law, check that the refit reproduces the published exponent, and test whether the gap between two exponents is distinguishable from zero. Across widely cited comparisons from eight released model suites, most claimed differences in scaling rate are statistical ties. The conclusion that the vanilla Transformer scales best among ten architectures does not hold against its close contenders, and nearly all pairs of pretraining corpora in the DataDecide suite differ in loss level but not in rate. The probability that the claimed winner would reverse, computed from the same test, separates these ties from large, well-established differences. A closed-form bound on the smallest detectable gap explains the pattern: the small sets of model sizes in common use cannot resolve gaps of the size being claimed. A survey of recent scaling papers further shows that most do not release the data needed to check their own comparisons. We close with simple reporting practices that make comparative scaling claims testable.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.