acceptodds
Under review as a conference paper at ICLR 2027

WHAT EMPIRICAL TRAINING DYNAMICS DO NOT DETERMINE: RELU HISTORIES AND ROBUST PREDICTION SEPARATION

Abstract

We study whether a ReLU network's empirical training record determines its breakpoint history, the order in which its breakpoints cross during training. The record includes training predictions, loss, the empirical neural tangent kernel, parameter and gradient norms, and the Hessian spectrum. For shallow one-dimensional networks with two training samples, we find an open set of initial parameters. From each point in this set, every stretchable complete reversal can occur from another initialization with the same record on a common finite gradient-flow interval. We also give a common gradient-descent schedule. A three-sample construction gives a prediction gap above 0.04 on a fixed interval after 16 steps at learning rate 0.005. This gap survives prescribed data perturbations and independent initialization and update errors. Further examples use two-dimensional data and two hidden layers, including a training activation change. These are selected constructions, not claims about typical initialization or generalization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.