Rounding in Preconditioner Space: Redesigning 4-bit AdamW Optimizer-State Quantization
Abstract
Quantizing AdamW's optimizer states reduces persistent storage, but quantization errors propagate through the moment recurrences and perturb subsequent adaptive updates. We redesign 4-bit optimizer-state quantization for AdamW from the perspective of rounding space: the coordinate in which a quantizer chooses between adjacent reconstruction levels. For the second-moment, a local analysis of the quantization cell adjacent to zero shows that small mean state error does not ensure a small mean error in the next-step preconditioner. A one-dimensional quadratic construction further shows qualitatively different optimization dynamics under state-space and preconditioner-space rounding. These results motivate Zero-Inclusive Preconditioner-space Stochastic Rounding (ZIP-SR), which retains zero in the second-moment codebook and computes stochastic-rounding probabilities in preconditioner space. As a complementary route, Zero-Excluding EDEN calibration (ZE-EDEN) uses a zero-excluding second-moment codebook and rescales the quantized second-moment block to mitigate the preconditioner distortion caused by the positive quantization floor. Both configurations use 4-bit NormalFloat (NF4) for the first moment, with targeted stochastic rounding of the LM-head first moment during the final 10% of training. Across GPT- and Llama-style pretraining experiments ranging from 130M to 2.7B parameters, both methods reduce TorchAO 4-bit AdamW's validation-loss gap to 32-bit AdamW at every evaluated scale, with the largest reported gap reduction reaching 70%. In full-parameter supervised fine-tuning, both recipes achieve lower validation loss than TorchAO while remaining close to 32-bit AdamW on downstream tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.