Rethinking Memory-Efficient Optimizers for LLM Pre-Training
Abstract
A growing family of memory-efficient optimizers reduces state storage through gradient projection, moment factorization, and shared statistics. Yet their practical value depends on when optimizer states constrain training and whether structural compression offers an advantage over simple state quantization. We examine these questions through a broad benchmark of memory-efficient methods and strong conventional optimizers under a common language-model pretraining setup. Our evaluation covers 257M and 500M dense models and extends to mixture-of-experts architectures, varying compute and state precision independently. We investigate how compression strength affects training quality and compare the approaches in terms of peak memory, validation loss at fixed token budgets, and time to quality. In the main 500M comparison, Muon with FP8 states achieves lower validation loss and lower peak training memory than the evaluated structural methods with FP32 states under both BF16 and FP8 computation. Quantizing the states of structural methods remains feasible, but yields limited further memory savings over quantized-state Muon without closing the quality gap. These findings support quantizing a strong optimizer as a practical starting point for memory-constrained pretraining and make competitive quantized baselines essential to assessing the benefits of structural compression.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.