acceptodds
Under review as a conference paper at ICLR 2027

How Model Growth, Boundary Operators, and Recursion Improve Scaling Laws

Abstract

Scaling laws predict how loss decreases as computation increases. Yet how architecture affects scaling laws is rarely examined. While a better constant brings scale-invariant compute-efficiency gains, a better exponent leads to power-law improvements in performance as computation increases. In this work, we comprehensively study how model growth, recursion, and boundary operators affect scaling coefficients. Specifically, we run more than 30 distinct scaling ladders over compute budgets of approximately to FLOPs. We find that better architectures can improve the *exponents*. In particular, our best architectural variant, *Untied-Grow*, achieves a compute-efficiency gain at FLOPs, and the trend extrapolates to FLOPs, where our model matches GPT-3 13B on CORE. We find that weight sharing is suboptimal in fresh-token training. However, recursion has a useful regularizing effect in data-constrained, multi-epoch training. In this regime, we find that it is compute-optimal to increase the number of loops with scale, outperforming model-size scaling even with tuned weight decay. To understand these gains, we examine computational depth and find that exponent improvements correlate with faster growth in computational depth.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.