acceptodds
Under review as a conference paper at ICLR 2027

Nesterov Acceleration in Local Step Methods

Abstract

Local-step methods reduce communication in distributed training by allowing workers to perform multiple optimizer steps between synchronizations. Prior work has shown that incorporating Nesterov acceleration can improve local-step training, but how it should interact with both periodic communication and updates within each local block remains less understood. We identify four constructions for such interaction, each restricted to communicating one model-size vector between workers per round. We analyze classical, modern, and primal formulations across these constructions, including generalized variants with independent control of smoothing and look-ahead, and establish when they are equivalent. Our language-modeling experiments identify two ways to improve DiLoCo-style training: tuning generalized retention and extrapolation coefficients, and reporting a smoothed iterate. They also identify a construction that performs well with frequent communication, and in what settings Nesterov acceleration does not help.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.