Biax: Fixed-Radius Matrix Optimization with Factored Preconditioning for Transformers
Abstract
Matrix optimizers for Transformers must reconcile spectral direction, anisotropic gradient scale, and parameter-norm control. We present Biax, which combines Muon²-Factorized-style temporal row/column profiling before polar extraction with Hyperball’s per-matrix fixed-initial-radius projection, adding (O(m+n)) persistent state. The profile selects a spectrally balanced direction, while sphere projection controls parameter scale. We show that the idealized update is a profiled-covector spectral linear oracle followed by exact sphere projection, separating anisotropic direction selection from parameter-scale control. We evaluate Biax in a paired five-tier, 2B-token language-model study spanning 10M to 1B parameters, with additional 500M and 1B experiments trained for 20B tokens. We also conduct an adaptive four-condition study at 10M, comparing full and row/column-factored profiles both with and without fixed-radius projection. Factoring lowers mean NLL in both comparisons, but with four seeds neither its overall effect nor its interaction with fixed-radius projection is statistically resolved. In the 2B-token panel, Biax lowers mean held-out NLL against its fixed-radius parent at every scale and wins all four paired seeds at 500M and 1B. Additional tangent-history and final-direction projection are inferior at both large 2B-token tiers. These results identify Biax as a simple, low-state optimizer and demonstrate why optimizer mechanisms must be evaluated against their nearest parents across scale.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.