Convergence of Mean-Field Transformers and Bilevel Training Dynamics in the Repulsive Regime
Abstract
We study token convergence and parameter training in continuous-depth, mean-field self-attention on the sphere, in the repulsive regime with parameters shared across depth. For fixed symmetric positive-definite interactions, we investigate a Polyak–ojasiewicz-type inequality relating energy dissipation to a power of the energy gap. Under uniform regularity and non-degeneracy bounds along the flow, we establish this inequality with exponent , yielding an energy-gap bound. We then study training as a bilevel optimization problem: the token distribution evolves toward equilibrium, while the parameters are adjusted to minimize a loss defined at that equilibrium. Tokens and parameters evolve simultaneously, without requiring equilibration between parameter updates. Under suitable regularity and gradient-approximation assumptions, a dissipation exponent determines a learning-rate schedule yielding an bound on the time-averaged squared training-loss gradient. In particular, quadratic dissipation gives an training bound. Numerical experiments support the predicted convergence rates.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.