Unifying Local Communication and Local Updates for LLM Pretraining
Abstract
Several methods reduce the communication cost of distributed LLM pre-training by performing multiple local optimizer steps between global aggregations. However, these methods still rely on global collectives, whose duration is limited by the available bandwidth and the slowest worker. We introduce GASLoC, a decentralized method in which each worker performs one or more local optimizer steps, exchanges the resulting parameters with one or two randomly selected peers, and applies an outer optimizer to the exchanged update. This structure avoids global collectives and allows the number of local steps to vary across workers. We establish convergence under different local step counts and show that outer momentum can accelerate consensus. Empirically, at token budgets selected according to compute optimal scaling, the strongest GASLoC variant achieves lower validation loss than the evaluated decentralized baselines when communicating after every local step. With multiple local steps between communications, it remains competitive with DiLoCo while avoiding global collectives. Under heterogeneous bandwidth, adapting the number of local steps across workers reduces end-to-end training time relative to DiLoCo.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.