Mean-Field Langevin Dynamics for Two-Layer ReLU Networks
Abstract
Mean-field Langevin dynamics provide a framework for analyzing convergence of wide two-layer neural networks while capturing feature learning. But standard uniform log-Sobolev arguments do not apply to ReLU activations with weak weight decay, where the associated Gibbs potentials may be non-coercive. For square loss, under boundedness, regularity of the activation indicator, and a mild initialization condition, we prove existence of global weak solutions and establish exponential convergence of the free energy to its global minimum for every positive weight-decay and temperature parameter. The proof uses a positive coercivity threshold to separate two regimes. In the low-coercivity regime, activation regularity and two-homogeneity yield a uniform lower bound on the dissipation; in the coercive regime, radial rescaling yields a log-Sobolev inequality. Together, these estimates imply a Polyak–Lojasiewicz inequality on every free-energy sublevel. For fixed input dimension and other problem parameters, the resulting convergence rate deteriorates at most polynomially in the inverse temperature.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.