acceptodds
Under review as a conference paper at ICLR 2027

Decided Early, Seen Late: The Committor of Neural Network Training

Abstract

Training a neural network compiles one global number, the loss, into millions of weights, and one number cannot pin them all down: two runs of one recipe can fit their data equally well and still differ in whether they generalize or in which attention head carries a capability such as in-context copying. We ask which source of chance makes that choice, when it is made, and whether anything in the network shows it while it is being made. To answer all three, we copy a checkpoint into replicas that differ only in the randomness that follows, as chemical physics does to estimate a committor, and divide the variance of an outcome into the part chance has yet to decide, the part already decided that a description of the network can read, and the part it cannot. The instrument recovers what was known by other means, such as early stability to training noise and the equal effect of every source of randomness on a ResNet's errors. It also shows that (1) equally good runs differ along only a few directions, along which chance moves a run, (2) which source of chance decides depends on the recipe, (3) outcomes which are decided at all are decided early, at a median of 7.5% of training in our recipes and long before they show in the run, and (4) only a task readout or a projection of the network's outputs reads the decision as it is made, whereas the diagnostics commonly used to watch training read it thousands of steps later yet pass the usual within-run test of a precursor. The instrument needs nothing more than the ability to restart a run from a checkpoint, so one can use it to date when an outcome of their own recipe is decided, to check whether a proposed diagnostic reads that decision or only its aftermath, and to forecast whether a run will generalize by letting a few replicas run briefly.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.