Nesterov Acceleration in Local Step Methods
Abstract
Local-step methods reduce communication in distributed training by allowing workers to perform multiple optimizer steps between synchronizations. Prior work has shown that incorporating Nesterov acceleration can improve local-step training, but how it should interact with both periodic communication and updates within each local block remains less understood. We identify four constructions for such interaction, each restricted to communicating one model-size vector between workers per round. We analyze classical, modern, and primal formulations across these constructions, including generalized variants with independent control of smoothing and look-ahead, and establish when they are equivalent. Our language-modeling experiments identify two ways to improve DiLoCo-style training: tuning generalized retention and extrapolation coefficients, and reporting a smoothed iterate. They also identify a construction that performs well with frequent communication, and in what settings Nesterov acceleration does not help.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.