PACE: Parameter Averaging with Consensus Extrapolation for Distributed Training
Abstract
Existing methods for training language models across slow, heterogeneous, or unreliable networks utilize infrequent or compressed gradient exchanges to reduce communication traffic. However, this creates two problems. a) They still pause computation and b) work best only when every worker remains available. Asynchronous sparse parameter averaging avoids both constraints: workers exchange a random subset of model weights in the background, absent workers contribute nothing, and stale workers are naturally pulled back when they return. Its weakness is optimization quality. We show that the missing ingredient is momentum on the mean motion of the worker population. That motion is already present in each sparse average, so Pace recovers it as a shared consensus-velocity buffer without sending any additional data. Every worker reconstructs the same buffer and uses it in a Nesterov-style update. We also propose accumulate-to-match (ATM) for handling heterogeneous hardware. ATM lets each worker take an optimizer step only after processing the same number of samples, so slower workers become stale instead of providing noisy updates due to lower batch sizes. Sparse averaging handles this bounded staleness much better than such gradient noise. Together, these changes turn sparse parameter averaging from a weaker optimizer into the strongest communication-efficient method in our heterogeneous and unreliable settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.