Clustering Stability and Collapse in Deeper Networks
Abstract
We show that as neural networks grow deeper, their internal representations progressively lose the clustering structure of the input and that this degradation, measured by the Frobenius distance between projection operators, is tightly and consistently associated with a sharp drop in clustering accuracy across architectures and datasets. This geometric characterization leads directly to two principled strategies: a structure-aware regularizer that penalizes projection distance at each layer, and an orthogonal initialization strategy that provides a structurally favorable starting point for training. Across more than 60,000 trials spanning five prevalent architectures and eight real-world datasets, our regularizer delays collapse in performance by an average of 7 layers and sustains high clustering accuracy well past the depth at which all unregularized counterparts have collapsed entirely.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.