Activation tails select suboptimal profiles in expressive networks
Abstract
ReLU homogenization removes Gaussian activation tails, but those tails can still determine the leading parameter profile of a positive-risk training trajectory. For any width m ≥ 4 of the hidden layer and smoothing parameter 0 < τ ≤ 1/5, we assume a bias-free one-hidden layer neural network with Gaussian-smoothed ReLU activations. All the weights and readouts are trained via population logistic gradient flow on the eight-point binary classification task. A nonempty open set of sufficiently large finite initializations gives rise to trajectories that misclassify two points forever. Both the smooth network and its width-m homogeneous ReLU version are able to classify all eight points. However, the loss for the trajectories of the smooth network tends to a positive, non-optimal population loss. The parameter vector, normalized by √log(t + t₀), converges to one point on an explicitly constructed sphere of local constrained norm minimizers. Four constraints correspond to the margins obtained from the points that were correctly classified, while four constraints are related to the activation tails of the two points that were misclassified. The absence of the latter four constraints yields a different norm-minimizing parameter profile. Under the same Euclidean parameter metric and common layer learning rate, these Gaussian parameter profiles are incompatible with the neuronwise balance invariant of homogeneous ReLU gradient flow. Different smoothing parameters yield the same ReLU homogenization but different leading log(t + t₀)-normalized logits at a specified off-support probe, and different limiting ratios to a reference training logit. The bad basin and profile convergence property are preserved under sufficiently small perturbations of each input coordinate independently and arbitrary positive probability weights.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.