acceptodds
Under review as a conference paper at ICLR 2027

ReAdam: Reformulating Adam for Low-Bit Quantization

Abstract

Adaptive optimizers like Adam are very effective for model training, adaptively adjusting the learning rate for each learnable parameter using first- and second-order gradient statistics. However, storing these moments substantially increases Adam’s memory footprint. Existing memory-efficient optimizers aim at preserving Adam’s optimizer state representation and quantizing the first- and second-order moments independently. However, these raw moments are unbounded and prone to outliers, thus requiring more elaborate and memory-intensive quantization schemes. We pose a different question: instead of improving how Adam’s existing optimizer state can be quantized, can the optimizer state itself be reparameterized to be more amenable to quantization? To this end, we derive a new scale-normalized reformulation of Adam in READAM by separating the optimizer state into (i) a gradient scale and (ii) a normalized quantity corresponding directly to Adam’s effective update. We establish that the square root of the second moment is a particularly effective choice for the scale. READAM is an equivalent reformulation of Adam using these two new optimizer coordinates, that are much more compression-friendly compared to the raw moments of Adam. READAM performs on par with or better than existing low-bit optimizers across multiple benchmarks and modalities (text, speech and vision), while also reducing optimizer-state memory usage by 6% (relative).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.