Understanding Implicit Averaging in Schedule-Free Stochastic Gradient Descent
Abstract
Schedule-Free SGD (SF-SGD) eliminates the need for predefined learning-rate schedules by combining momentum updates with implicit averaging. A distinctive feature of SF-SGD is that it generates two iterates: the gradient-location point , where the stochastic gradient is evaluated, and the evaluation point , which implicitly aggregates information from past gradient-location iterates. Despite the practical importance of this averaging mechanism, its effect on the relationship between these two points remains insufficiently understood. We study SF-SGD under nonconvex objectives with constant averaging and establish convergence guarantees for both and , together with an upper bound on their gradient-norm gap. In particular, we show that the upper bound at can be smaller than that at , providing theoretical evidence that implicit weighted averaging can improve the stationarity of the evaluation point. We also derive sufficient conditions characterizing the admissible ranges of the learning rate, momentum parameter, and averaging coefficient and extend our analysis to mini-batch stochastic gradients without assuming bounded gradients. Numerical experiments show that constant averaging can outperform uniform averaging in terms of gradient norms, gradient-norm gaps, and test accuracy while being more robust to the momentum parameter. We also observe advantages of increasing over fixed batch sizes under both matched training-sample budgets and equal numbers of optimization steps. The code is available at https://anonymous.4open.science/r/schedule_free_convergence-CBAD, which also includes additional experimental results.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.