acceptodds
Under review as a conference paper at ICLR 2027

Precision-Aware Optimal Newton–Schulz Iterations for the Matrix Sign

Abstract

The Muon optimizer orthogonalizes each gradient matrix by running a short quintic Newton–Schulz (NS) iteration that drives its singular values to one—an approximate matrix sign function, msign(M) = UV^T. Recent work (Polar Express) shows that the per-step coefficients of this iteration should not be hand-tuned but solved for: a per-step minimax (Remez) fit yields provably fastest convergence in exact arithmetic. But Muon actually runs the iteration in bfloat16, and no prior analysis characterizes these learned iterations under low precision. We give a precision-aware characterization of the coefficient design space, deliberately scoped to the idealized scalar reduction of the iteration—the singular-value band [1/κ, 1] with per-scalar rounding—and bounded honestly against real gradients and real training. We study the sign iteration as an optimal–structured–hybrid family—P1, a free per-step minimax fit (Polar Express); P2, the same minimax under a boundedness constraint; and P3, a hybrid that runs free steps early and bounded steps late—and measure error-versus-step curves in fp32 and bf16 across κ ∈ 10, 30, 100. Three findings. (i) On the idealized band, the fp32 per-step-optimal P1 crushes fixed Muon at a matched budget (1e-8 vs. 0.318 at κ = 10; 7.8e-6 vs. 0.318 at κ = 100). (ii) On that same idealized band, bf16 is fragile at high conditioning: at κ = 100 the aggressive P1 coefficients make the bf16 scalar iteration diverge (terminal 4×10^18; P2 overflows to +∞), while the hybrid P3 stays finite (bf16 floor 0.141, beating fixed’s 0.328); we give a first-order model of this floor. (iii) This bf16-specificity is a feature of the idealized band, not of real gradients. Running the true matrix iteration on real 124M gradient matrices (fp32-accumulate tensor cores), the optimal P1 coefficients are fragile in both precisions—blowing up on 8/8 matrices in bf16 and 7/8 in fp32, so the clean bf16-only divergence appears for just 1/8—while fixed Muon stays stable (residual ≈ 0.32). And the deployed iteration is bf16-safe end-to-end: a nanoGPT-124M premise test trains equally well in bf16 and fp32 (a null, −0.038 nats). Our contribution is thus a precision-aware map of the coefficient design space in the idealized setting, honestly bounded: the optimal-coefficient instability does not clearly transfer to real gradients bf16-specifically, and today’s Muon is already bf16-safe—not a fix for a real deployment problem. Extending precision-aware optimal iterations to general spectral functions is left to future work.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.