SGD Selects Initialization-Dependent Representation in Overparameterized ReLU Networks
Abstract
Overparameterized neural networks can represent the same function through many different hidden configurations, raising the question of which representation is favored by a given training algorithm. We study how stochastic gradient descent (SGD), jointly training both layers, resolves this ambiguity in finite-width, bias-free, two-layer ReLU networks trained to learn a linear teacher under Gaussian inputs. At width two, exact fitting leaves no choice: the neurons must align with the teacher directions. At every width , however, infinitely many non-aligned exact fits are possible. Despite this ambiguity, we prove that joint batch-size-one SGD from sufficiently small initialization drives every neuron arbitrarily close to one of the two teacher directions while reaching arbitrarily small population error in finite time. This occurs without explicit regularization, layerwise training, gradient clipping, or a learning-rate schedule. The dynamics additionally select an initialization-dependent allocation of the fitted function across neurons, providing a finer characterization of the representation selected by SGD. Numerical experiments reproduce the predicted alignment, fitting, and initialization-dependent allocation. Our results identify sufficient conditions under which standard SGD exhibits a strong implicit bias toward a highly structured internal representation, even when overparameterization creates infinitely many qualitatively different exact representations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.