Removing spurious minima for planar features by skip connections
Abstract
Understanding loss landscapes is central to explaining neural-network training, yet their structure remains only partially understood even in simple models. We study the Gaussian population loss of shallow, bias-free ReLU networks in the teacher–student setting, which provides a simple model for studying essential aspects such as feature learning and overparameterization. For teachers with positive output weights and planar features, we show that including a learned linear skip removes all spurious local minima with nonnegative student output weights, provided that the student is at least as wide as the teacher. In contrast, without the skip, we construct a fixed teacher with positive output weights and only three hidden neurons in input dimension two whose spurious local minima persist at every student width at least three. Thus, a learned linear skip can remove spurious minima that persist under arbitrary overparameterization. We further show that a positive output weight student always learns the subspace spanned by the teacher features: student features at local minima with nonnegative student output weights lie in the span of the teacher features. For ReLU networks in two dimensions, even heavily overparameterized students have effective width controlled by the teacher width: every critical point with positive student output weights has at most twice as many distinct student feature directions as teacher neurons. Finally, we transfer the benignity result to empirical minima over parameter balls of any prescribed radius, with the required sampling accuracy depending on that radius.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.