Parameter Conventions Confound Compute-Optimal Scaling Exponents of Dense Language Models
Abstract
Compute-optimal scaling exponents are routinely compared across labs as if they measured the same quantity. We audit every public language-model scaling sweep we could find whose metadata determine, run by run, exactly how many parameters sit in vocabulary-sized matrices: 18 sweeps (1,673 runs) from six independent research groups. Changing only whether those matrices are counted raises the fitted data-model exponent, r, in every sweep and every group, by 0.54 on average across groups in held-out fits, comparable to the entire Kaplan-Chinchilla gap, under both separable and non-separable loss forms. Public data do not consistently settle which count is right: small-to-large extrapolation leans toward transformer-only counting in most groups but not significantly, holding out the largest models splits the groups three to three, and only one lab's frontier holdout favors one count clearly. The convention changes each exponent, but we find no evidence that it explains why labs disagree: harmonizing every sweep to one convention does not detectably reduce the cross-lab spread relative to the mean exponent. The two conventions nonetheless prescribe different frontier allocations, a median 2.3 times difference in tokens per parameter at 100 times the fitted compute. We release exact reconstructions, a convention registry, and a reporting standard.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.