We should Understand Symmetry Breaking in Sparse Training to Unmask the Role of Overparameterization
Abstract
Overparameterization is a cornerstone of deep learning. When combined with weight-space symmetries, it leads to many equivalent solutions in the loss landscape, making it easier for an optimizer to find a good solution. However, breaking these symmetries — as in Sparse Neural Networks (SNNs) — reveals fundamental limitations in our training algorithms, present in both sparse and dense settings. Although in the literature sparse training research has mostly been motivated by making training efficient, in this position paper, we argue that sparse training with a fixed mask provides a valuable testbed for studying symmetry, optimization, and their interplay in a principled manner. We explore novel examples of how breaking permutation symmetry and rescaling invariance jointly affect optimization, and can inspire algorithmic advancements. Sparse training research can provide valuable insights into the role of symmetries and overparameterization in training in general, not limited to sparse training. This position paper argues that understanding how symmetry breaking can improve or hinder training could uncover deeper theoretical insights and inform the design of more robust optimizers and architectures that respect symmetries.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.