acceptodds
Under review as a conference paper at ICLR 2027

MOSES: Geometric Preconditioning and Hybrid Quantization for Memory-Efficient LLM Optimization

Abstract

Optimizer states are the largest static memory cost in large model training. We introduce MOSES, a memory-efficient optimizer that addresses this challenge through two core innovations. First, Geometry-Aware Adaptive Preconditioning integrates gradient orthogonalization with factorized second-moment estimation, leveraging the geometry of the gradient space to provide both spatial smoothing and temporal adaptivity with negligible memory overhead. Second, Distribution-Aware Hybrid 1/4/8-bit Quantization adaptively allocates precision to post-preconditioning momentum based on the spectral properties of EMA updates, enabling aggressive compression while preserving optimization quality. Together, these innovations reduce optimizer state memory by 12.3× compared to Adam. Across model scales from 125M to 6.7B parameters, Moses achieves a compelling memory–performance trade-off that existing memory-efficient methods struggle to match, particularly excelling in the small-learning-rate regime that is characteristic of large-scale production training. Comprehensive evaluations across pretraining and downstream tasks show that Moses achieves competitive optimization quality while significantly improving memory efficiency, thereby reducing the hardware requirements for large language model training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.