Do Orthogonality Constraints Help Language-Model Training Beyond Sphere Optimization?
Abstract
A growing body of work reports benefits from orthogonality constraints in fine-tuning large language models (LLMs). These results raise an open question: do orthogonality constraints also help LLM pre-training? We conduct a systematic empirical study of decoder-only LLM pre-training with models ranging from 360M to 3.1B parameters, trained on nearly 200B tokens. To explore the potential of orthogonality in pre-training, we examine seventeen variants of orthogonality training: some constrain the weight matrix itself, while others use orthogonal factors to reparameterize the weight. Given that requiring a square weight matrix to be strictly orthogonal results in a loss of half of its parameter degrees of freedom, we also investigate block-wise orthogonality, where every rows of the weight matrix are required to form a row-orthonormal (Stiefel) matrix, with the block size serving as a hyperparameter. Among block sizes ranging from 1 to the shorter side of the matrix, no reliably outperforms , and large blocks are clearly worse. However, when , the orthogonality constraint degenerates into sphere optimization, meaning each row direction is optimized on its own unit sphere, which by itself already improves on tuned Muon optimizers. Our conclusion is that, across the realizations we tried, orthogonality constraints do not help language-model pre-training beyond sphere optimization. Since this appears to conflict with several prior results, we re-examine five prior methods in pre-training and fine-tuning that report orthogonality improvements. In each case the reported gain is attributable to numerical precision, training horizon, optimizer realization, or per-row norm control rather than to orthogonality between rows.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.