Hyperball Optimization for LLM Pretraining
Abstract
Matrix-based optimizers such as Muon can substantially speed up language model pretraining, but their gains over AdamW are observed to shrink as model size and data scale grow when using standard constant decoupled weight decay. We propose **Hyperball**, a simple optimizer wrapper that addresses this issue. The method is motivated by prior theory showing that weight decay drives the weight norm toward an equilibrium, eventually making the weights move only on a hypersphere of fixed radius. With Hyperball, we explicitly control this behavior: given a base optimizer such as Adam or Muon, Hyperball sets the Frobenius norms of weight matrices and their corresponding optimizer updates to fixed constants. On Qwen3 style models up to B parameters, Muon Hyperball achieves – token equivalent speedup over weight decay baselines. Hyperball also improves learning rate transfer across widths and depths compared to decoupled weight decay.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.