Phases of Cross-Entropy Training: Emergence and Implicit Bias in the Unconstrained Features Model
Abstract
A classical result for linear networks trained with mean-squared-error loss is that gradient flow learns the singular modes of the data sequentially, in order of importance. The precise mechanics of such sequential emergence under cross-entropy loss remain largely unknown. We study this question in a minimal nonconvex setting: a two-layer linear network with orthogonal inputs and step-imbalanced classes, equivalent to the unconstrained feature model used in neural collapse analyses. First, we characterize the regularization path in spectral coordinates, showing sequential mode emergence with activation thresholds that, unlike in MSE, need not coincide with the corresponding data singular values. As regularization vanishes, active singular values diverge, while their normalized values converge at a logarithmic rate and can approach their limits from either side. Second, for gradient flow under suitable spectral initialization, we prove majority-first emergence and show that the majority-to-minority mode ratio converges to that of the regularization-path limit. Our analysis relies on a novel imbalance-adapted Hadamard basis in which we show that softmax preserves a diagonal-plus-rank-one structure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.