Discrete Weight Updates for Extreme Low-Bit Training without Master Weights
Abstract
Memory consumption has become a major bottleneck for training and deploying increasingly large models. While extreme low-bit models with binary or ternary weights remarkably reduce the memory required for inference, their training still commonly relies on full-precision master weights that are optimized in a contin- uous space and quantized for the forward pass. Removing such master weights is challenging, as discrete weights can only change through transitions between a finite set of states and thus cannot directly represent the small continuous updates made by conventional optimizers. In this paper, we propose Discrete Updates for Extreme Low-bit Training (DUET), a scalable framework that directly trains ex- treme low-bit weights without maintaining master weights. We first derive a tran- sition budget from the magnitude of a continuous optimization step, specifying the expected number of discrete weight transitions. Given this transition budget, we select the weight transitions that maximize the alignment of the resulting discrete update with the continuous step. We further show that optimizer states used for these transition decisions can be significantly compressed while retaining training quality. Our experiments demonstrate that DUET maintains competitive training quality over long training horizons and on large models, with training memory reductions that increase with model size and reach 12.0× at 235B parameters compared to conventional training with master weights
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.