Predicting and Steering the Emergence of Capabilities in Neural Networks
Abstract
Neural networks often acquire capabilities abruptly during training, from in-context learning to arithmetic and multi-step reasoning. Scaling laws describe when a capability emerges on average, but not whether a particular run will acquire it, or whether that outcome can still be changed. Here, we introduce the training committor, the probability that a capability emerges within a fixed horizon from a given checkpoint, which we estimate by restarting training from the checkpoint many times. In small transformers and in a language model trained on natural text, we find that most runs are effectively decided before the capability appears, and that the committor predicts which runs will acquire it. Notably, on the controlled tasks, it is more accurate than predictors based on loss or weight statistics. Beyond prediction, the committor reveals where to intervene: a short burst of training data aligned with the capability can rescue nearly all of the runs it marks as unlikely to succeed. Since its branches can be trained under any intervention, the committor can also compare interventions on the same run and pick the most effective one. Overall, our results establish capability emergence as a predictable and steerable property of the training state.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.