acceptodds
Under review as a conference paper at ICLR 2027

The Stabilizing Effect of Gossip in Decentralized Language Model Pre-Training

Abstract

As language-model training scales beyond a single connected cluster, every-step global synchronization becomes increasingly expensive. Gossip-based decentralized training replaces global all-reduce with local updates and peer-to-peer parameter averaging. We investigate how this change affects training dynamics, beyond its communication benefits. Across the reported model families and optimizer settings, gossip-based decentralized training maintains comparatively stable validation-loss trajectories in large-learning-rate regimes where centralized training becomes highly oscillatory or moves toward high loss. To understand this phenomenon, we compare the training dynamics of decentralized and centralized SGD in deep linear networks. We prove that gossip SGD can achieve strictly lower expected full-Hessian sharpness at the worker-averaged iterate than centralized SGD. Our analysis further establishes conditions under which gossip SGD improves stability to directional parameter perturbations at large learning rates, quantified by reduced multistep sensitivity. This provides a theoretical counterpart to the stability advantage observed in our language-model experiments.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.