acceptodds
Under review as a conference paper at ICLR 2027

The Accuracy Well: Nonmonotone Accuracy Dynamics at Vanishing Initialization

Abstract

We study classification dynamics under small initialization and identify a pronounced non-monotone behavior of accuracy that we call the *accuracy well*: accuracy rises rapidly early in training, subsequently drops, and only much later recovers. We observe this phenomenon in deep linear networks, fully connected ReLU networks on MNIST, deep convolutional networks on CIFAR-10, and controlled synthetic tasks, across different losses and optimizers and in both training and test accuracy. We propose a three-stage explanation in terms of target-driven organization, a degeneracy bottleneck, and subsequent recruitment and recovery. We develop a rigorous theory of these dynamics for shallow and deep linear networks, showing in particular that high accuracy can emerge while the predictor and the change in loss remain asymptotically vanishing, before continued optimization causes accuracy to deteriorate. For a simple synthetic classification task, we prove an extreme instance in which accuracy asymptotically rises to , collapses to , and later recovers permanently to . Our results show that classification accuracy can reveal aspects of learning dynamics that are not captured by the training loss alone.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.