acceptodds
Under review as a conference paper at ICLR 2027

Hyperball Optimization for LLM Pretraining

Abstract

Matrix-based optimizers such as Muon can substantially speed up language model pretraining, but their gains over AdamW are observed to shrink as model size and data scale grow when using standard constant decoupled weight decay. We propose **Hyperball**, a simple optimizer wrapper that addresses this issue. The method is motivated by prior theory showing that weight decay drives the weight norm toward an equilibrium, eventually making the weights move only on a hypersphere of fixed radius. With Hyperball, we explicitly control this behavior: given a base optimizer such as Adam or Muon, Hyperball sets the Frobenius norms of weight matrices and their corresponding optimizer updates to fixed constants. On Qwen3 style models up to B parameters, Muon Hyperball achieves – token equivalent speedup over weight decay baselines. Hyperball also improves learning rate transfer across widths and depths compared to decoupled weight decay.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.