acceptodds
Under review as a conference paper at ICLR 2027

SGD Selects Initialization-Dependent Representation in Overparameterized ReLU Networks

Abstract

Overparameterized neural networks can represent the same function through many different hidden configurations, raising the question of which representation is favored by a given training algorithm. We study how stochastic gradient descent (SGD), jointly training both layers, resolves this ambiguity in finite-width, bias-free, two-layer ReLU networks trained to learn a linear teacher under Gaussian inputs. At width two, exact fitting leaves no choice: the neurons must align with the teacher directions. At every width , however, infinitely many non-aligned exact fits are possible. Despite this ambiguity, we prove that joint batch-size-one SGD from sufficiently small initialization drives every neuron arbitrarily close to one of the two teacher directions while reaching arbitrarily small population error in finite time. This occurs without explicit regularization, layerwise training, gradient clipping, or a learning-rate schedule. The dynamics additionally select an initialization-dependent allocation of the fitted function across neurons, providing a finer characterization of the representation selected by SGD. Numerical experiments reproduce the predicted alignment, fitting, and initialization-dependent allocation. Our results identify sufficient conditions under which standard SGD exhibits a strong implicit bias toward a highly structured internal representation, even when overparameterization creates infinitely many qualitatively different exact representations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.