Non-Convex Online SGD with Dependent Data: Last-Iterate Convergence and Temporal Preconditioning
Abstract
Stochastic gradient descent (SGD) is usually analyzed under the assumption that data samples are independent and identically distributed. In practice, however, samples collected over time are often correlated, which can bias stochastic gradients and slow down the training process. We study online SGD for non-convex optimization problems when the data follow a temporally dependent -mixing process. Our analysis explains how the sampling strategy, the rate at which data correlations decay, and the optimization dynamics together determine last-iterate convergence. In particular, we investigate how subsampling and mini-batch sampling methods mitigate the effects of data dependence and how their effectiveness depends on the mixing rate. We also propose a preconditioner that uses online estimates of temporal dependence to adjust SGD updates. We study this approach in controlled settings with known additive noise. The proposed preconditioner complements our convergence analysis of sampling strategies by addressing temporal dependence through the SGD updates.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.