acceptodds
Under review as a conference paper at ICLR 2027

Revisiting Data Sharding for Distributed Training with Periodic Synchronization

Abstract

In datacenter distributed training, periodic model synchronization can mitigate communication bottlenecks and accelerate training. However, longer synchronization intervals exacerbate convergence slowdowns induced by data heterogeneity across nodes. To improve convergence efficiency without increasing model synchronization, we revisit data sharding and propose a weighted stratified sharding algorithm. Our algorithm exploits objective values observed during training to construct weighted shards that reduce objective discrepancies across nodes. We establish theoretical guarantees for the induced weighted objectives and derive a convergence bound in terms of the objective approximation error. The experiments span various distributed setups and training tasks, from image classification to math continual pretraining. Compared with widely used random partitioning, our sharding algorithm introduces negligible runtime overhead and converges up to 1.45 times faster in training steps. It achieves both lower training loss and improved test accuracy, and these gains persist even when using nearly 80% fewer synchronization rounds.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.