The Stabilizing Effect of Gossip in Decentralized Language Model Pre-Training
Abstract
As language-model training scales beyond a single connected cluster, every-step global synchronization becomes increasingly expensive. Gossip-based decentralized training replaces global all-reduce with local updates and peer-to-peer parameter averaging. We investigate how this change affects training dynamics, beyond its communication benefits. Across the reported model families and optimizer settings, gossip-based decentralized training maintains comparatively stable validation-loss trajectories in large-learning-rate regimes where centralized training becomes highly oscillatory or moves toward high loss. To understand this phenomenon, we compare the training dynamics of decentralized and centralized SGD in deep linear networks. We prove that gossip SGD can achieve strictly lower expected full-Hessian sharpness at the worker-averaged iterate than centralized SGD. Our analysis further establishes conditions under which gossip SGD improves stability to directional parameter perturbations at large learning rates, quantified by reduced multistep sensitivity. This provides a theoretical counterpart to the stability advantage observed in our language-model experiments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.