Weight Decay and Neuron Condensation: A Three-Stage Analysis of Two-Layer ReLU Networks
Abstract
Weight decay is widely used as a regularization technique in neural network training, yet its role in neuron condensation (parameter direction alignment) remains unclear. Starting from a parameter initialization in the neural tangent kernel regime, we characterize training dynamics under weight decay through three stages: rapid fitting, amplitude compression, and neuron condensation. Using a two-layer ReLU network, we analyze a residual correlation field that governs both neuron amplitudes and directions. During rapid fitting, the residual approaches a quasi-static equilibrium maintained by weight decay while the tangent kernel remains nearly unchanged. In the early stage of amplitude compression, kernel decay amplifies the residual correlation field, whose isolated attracting extrema provide candidates for condensation directions. As neuron amplitudes stabilize, we bound the drift of attracting extrema and demonstrate contraction of neuron directions around them, leading to neuron condensation. This staged analysis provides a dynamical understanding of how weight decay promotes a condensed representation, beyond reducing parameter norms.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.