Symmetry and Optimization in Overparameterized Networks
Abstract
Overparameterization is central to the success of deep learning, yet the mechanisms by which it improves optimization remain incompletely understood. In hopes of better understanding how the loss landscape transforms as we add neurons, we study neural networks of different width that represent the same function. Building on prior work, we characterize function-preserving transformations across network width in terms of families of matrices. These width symmetries allow for the study of a hierarchy of solutions across neural network width. We then use this construction to derive consequences for loss landscape geometry and optimization. First, we prove that these symmetries act as a form of diagonal preconditioning on the Hessian, enabling the existence of better-conditioned minima within each class of functionally identical solutions. Second, we find that overparameterization increases the probability mass of global minima near typical initializations, making these favourable solutions more reachable. These results offer a potential link between loss landscape geometry and simplicity bias. Empirically, we observe wider networks have lower top eigenvalues, smaller condition numbers and faster convergence, supporting our analysis. Together, these results describe width growth as a geometric transformation of the loss landscape.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.