acceptodds
Under review as a conference paper at ICLR 2027

SPARQ: A Subspace-Preserving Mixed Precision LMO Optimizer

Abstract

Optimizer-state memory remains a major bottleneck in large language model training. Existing low-bit state-compression methods do not explicitly account for the geometric sensitivity of linear-minimization-oracle (LMO) updates. This limitation is particularly important for Scion, whose spectral updates depend on the singular subspaces of persistent exponential-moving-average (EMA) states. We study Subspace-Preserving Adaptive Representation Quantization (SPARQ), a Scion-oriented mixed-precision state-compression method. SPARQ stores a more accurate low-rank component of each matrix-shaped EMA state and compresses the residual more aggressively, while applying the spectral LMO to matrix-shaped parameters and the root-mean-square (RMS) LMO to one-dimensional parameters. Our analysis shows that bounded reconstruction error preserves the decaying terms of the full-precision convergence bound while introducing an explicit error floor. Experiments with models up to 3B parameters show that SPARQ remains close to full-precision Scion, improves over uniform 4-bit integer (INT4) state storage, and reduces Scion optimizer-state memory by up to . By reducing optimizer-state memory while preserving training quality, SPARQ can facilitate the training of larger models under memory-constrained settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.