SPARQ: A Subspace-Preserving Mixed Precision LMO Optimizer
Abstract
Optimizer-state memory remains a major bottleneck in large language model training. Existing low-bit state-compression methods do not explicitly account for the geometric sensitivity of linear-minimization-oracle (LMO) updates. This limitation is particularly important for Scion, whose spectral updates depend on the singular subspaces of persistent exponential-moving-average (EMA) states. We study Subspace-Preserving Adaptive Representation Quantization (SPARQ), a Scion-oriented mixed-precision state-compression method. SPARQ stores a more accurate low-rank component of each matrix-shaped EMA state and compresses the residual more aggressively, while applying the spectral LMO to matrix-shaped parameters and the root-mean-square (RMS) LMO to one-dimensional parameters. Our analysis shows that bounded reconstruction error preserves the decaying terms of the full-precision convergence bound while introducing an explicit error floor. Experiments with models up to 3B parameters show that SPARQ remains close to full-precision Scion, improves over uniform 4-bit integer (INT4) state storage, and reduces Scion optimizer-state memory by up to . By reducing optimizer-state memory while preserving training quality, SPARQ can facilitate the training of larger models under memory-constrained settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.