The Rate Measure: How Far Linearisation Governs Training, and Why That Horizon Cannot Be Read
Abstract
The loss curve of a network trained by gradient descent at squared error can be predicted before it trains. The linearised loss is a Laplace transform of a measure carried by the output-space Gauss–Newton operator , and one Krylov pass before training returns it. We measure the horizon, how long real networks follow this forecast, and find it set by how far moves during training. From M to B parameters, standard-parameterised transformers follow the forecast through the whole window we score from M parameters up. Under maximal-update parameterisation the same networks leave it early. At one size and one initial operator, scaling the targets alone moves the horizon from the whole run to under two hundred steps. At a fixed example count, across size, parameterisation and target scale, runs whose moves further leave the forecast earlier. A tenth of the way into training, this motion flags the runs whose forecast will fail. On MNIST classifiers its relative size does, and every classifier it certifies holds. On transformers to B its growth along the residual does, withholding all but four of the runs that fail. Before the first step, the forecast tracks the measured rate curve of standard-parameterised transformers more closely, in median over log time, than every scaling law we fit. It matches the best predictor that sees none of the run, with no fitted parameter. We prove that in the worst case over networks sharing the linearisation's inputs, no function of them predicts the horizon, and that at cross-entropy the rate is asymptotically model-free.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.