acceptodds
Under review as a conference paper at ICLR 2027

AdasignSGD: Turning 1-Bit Worker Votes into Adaptive Updates

Abstract

Momentum-sign methods reduce synchronization traffic in distributed AI training by transmitting each worker’s local momentum with one bit per coordinate. However, these methods discard local momentum magnitudes and often fail to match the optimization quality of full-precision methods. We introduce AdasignSGD, which uses the majority sign to determine the update direction and the vote margin to set the update magnitude, taking larger steps on coordinates with stronger worker agreement. In C4 language-model pretraining, AdasignSGD outperforms all evaluated 1-bit worker-sign baselines in validation perplexity across 13 model and training-budget settings from 60M to 1B parameters. Its gap to dense AdamW narrows with model scale. At 1B, AdasignSGD closely matches AdamW in validation perplexity while delivering a complete-step speedup over fused AdamW on 32 H200 GPUs. AdasignSGD uses one model-sized momentum state while incurring only 14.1% of FP32 dense-DDP model-sized traffic at 128 workers. On 128 P100 GPUs, AdasignSGD delivers a complete-step speedup over AdamW for the 350M model. Together, these results show that one-bit worker votes can be turned into adaptive updates that combine strong optimization quality with a smaller memory footprint and faster distributed training steps at scale.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.