acceptodds
Under review as a conference paper at ICLR 2027

When Does Scale Trade for Time?

Abstract

Can a larger model trained briefly behave like a smaller model trained for longer? We study this question, which we call scale-time equivalence, in a deliberately small and controlled setting where it can be derived, measured, and falsified. Theoretically, we analyze a random subspace model, in which the trained network occupies a random -dimensional subspace of a larger nonlinear model. This model contains Gaussian linear random-feature models, a workhorse of scaling theory, as the special case of a quadratic loss, but, unlike them, allows the network output to depend nonlinearly on the trainable parameters. We prove that under fixed-learning-rate gradient flow, the function learned by a model of scale after time is approximated by a -independent trajectory evaluated at , with an error that decays as for any fixed value of . The bound grows exponentially only in the effective training progress , not along the scale-time tradeoff itself, so it guarantees equivalence between models of very different scales as long as both are large and their shared value of is moderate. Empirically, in MLPs and CNNs trained with fixed-learning-rate SGD on MNIST, SVHN, and CIFAR-10, the number of epochs needed to generalize follows an approximate power law in model scale, with a common exponent across datasets, architectures, and training-set sizes, and small models trained for long predict large models trained briefly. With Adam, which lies outside the theorem's assumptions, the equivalence visibly breaks. Finally, we conjecture, and give preliminary evidence, that where scale-time equivalence holds, parameter-wise and time-wise double descent share a mechanism: the relative rates at which signal and noise are acquired.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.