From Learning Rates to Rotational Schedules: Finite-Horizon Theory on the Sphere
Abstract
How should step sizes be scheduled across a finite training horizon? Cosine and warmup-stable-decay schedules are widely adopted in large-scale pretraining, yet existing convergence theory for such schedules is largely Euclidean and does not capture the rotational dynamics of normalized neural networks. In this work, we formulate schedule design directly in the space of rotational motion. Working within a geodesically convex spherical setting, we derive a last-iterate convergence guarantee for arbitrary rotational schedules, i.e. sequences of angular step sizes, that separates the sum of the scheduled step sizes from their distribution over time. We recover the same finite-horizon schedule-dependency as in Euclidean theory, and we solve it globally: the bound-optimal schedule is unique, and approaches linear decay as the horizon grows. Finally, we connect this result to Hyperball, where fixed parameter norm and normalized updates make the learning rate act directly as the rotational schedule. We find that our theoretical rotational bound predicts schedule performance in practice, evaluated on language model pretraining runs optimized with Hyperball. Consistent with the bound prediction, we find that cosine and linear schedules achieve comparable performance, both ahead of warmup-stable-decay and constant schedules, thus providing practical insight into the schedule design of Hyperball-style optimizers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.