acceptodds
Under review as a conference paper at ICLR 2027

Hyperball may not be a free lunch

Abstract

Norm-constrained optimizers perform well in large language model pretraining, but the source of their gains and their relationship with learning-rate scheduling remain unclear. We study their mechanisms and learning-rate schedules from a scale-invariance perspective. First, we derive the relationship between the scale and angular effective learning rates in preconditioned optimizers. Numerical analyses show that radial updates have a limited direct effect on one-step angular displacement. By adjusting the nominal learning rate to match the angular effective learning rate, we can largely reproduce the phase-dependent loss trajectories of Hyperball and non-Hyperball optimizers in both directions. Scale effective learning-rate alignment also yields training outcomes close to those of angular alignment. Theoretically, we establish a local dynamic regret bound under strict scale invariance, local geodesic convexity, and bounded directional gradients. Combining this analysis with a quotient-space optimization perspective, we propose Effective Learning Rate Stable–Decay (ELR-SD), a heuristic schedule that holds the scale effective learning rate constant before linearly decaying it once the gradient norm before preconditioning stabilizes. Experiments on dense, mixture-of-experts (MoE), and long-horizon pretraining tasks show that ELR-SD improves validation loss. These results further suggest that the effectiveness of long-horizon linear decay in Hyperball is linked to its control of angular step sizes, motivating the use of step-size design principles from convex optimization in large language model pretraining.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.