MOSES: Geometric Preconditioning and Hybrid Quantization for Memory-Efficient LLM Optimization
Abstract
Optimizer states are the largest static memory cost in large model training. We introduce MOSES, a memory-efficient optimizer that addresses this challenge through two core innovations. First, Geometry-Aware Adaptive Preconditioning integrates gradient orthogonalization with factorized second-moment estimation, leveraging the geometry of the gradient space to provide both spatial smoothing and temporal adaptivity with negligible memory overhead. Second, Distribution-Aware Hybrid 1/4/8-bit Quantization adaptively allocates precision to post-preconditioning momentum based on the spectral properties of EMA updates, enabling aggressive compression while preserving optimization quality. Together, these innovations reduce optimizer state memory by 12.3× compared to Adam. Across model scales from 125M to 6.7B parameters, Moses achieves a compelling memory–performance trade-off that existing memory-efficient methods struggle to match, particularly excelling in the small-learning-rate regime that is characteristic of large-scale production training. Comprehensive evaluations across pretraining and downstream tasks show that Moses achieves competitive optimization quality while significantly improving memory efficiency, thereby reducing the hardware requirements for large language model training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.