acceptodds
Under review as a conference paper at ICLR 2027

Convergence of Mean-Field Transformers and Bilevel Training Dynamics in the Repulsive Regime

Abstract

We study token convergence and parameter training in continuous-depth, mean-field self-attention on the sphere, in the repulsive regime with parameters shared across depth. For fixed symmetric positive-definite interactions, we investigate a Polyak–ojasiewicz-type inequality relating energy dissipation to a power of the energy gap. Under uniform regularity and non-degeneracy bounds along the flow, we establish this inequality with exponent , yielding an energy-gap bound. We then study training as a bilevel optimization problem: the token distribution evolves toward equilibrium, while the parameters are adjusted to minimize a loss defined at that equilibrium. Tokens and parameters evolve simultaneously, without requiring equilibration between parameter updates. Under suitable regularity and gradient-approximation assumptions, a dissipation exponent determines a learning-rate schedule yielding an bound on the time-averaged squared training-loss gradient. In particular, quadratic dissipation gives an training bound. Numerical experiments support the predicted convergence rates.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.