acceptodds
Under review as a conference paper at ICLR 2027

Implicit Sparse Regularization in Momentum Gradient Descent: A Shifted Early Stopping Window Induced by Momentum and Depth

Abstract

In high-dimensional sparse regression, optimization algorithms are known to induce implicit sparsity through the interplay of overparameterization and early stopping. However, existing theory largely focuses on (stochastic) gradient descent (GD), leaving momentum-based methods — despite their widespread empirical success — poorly understood. This paper provides a theoretical characterization of implicit sparse regularization in momentum gradient descent (MGD). We show that for depth- diagonal networks, MGD with early stopping achieves minimax optimal sparse recovery with sufficiently small initialization and learning rate. Moreover, increasing depth significantly enlarges both the admissible initialization scale and the effective early stopping window, making implicit sparsity more likely to emerge. More importantly, we uncover a distinct temporal effect of momentum: a larger momentum shortens the early stopping window, but it shifts the entire window toward earlier iterations. As a result, MGD induces sparsity earlier than GD. Experimental results strongly support our theoretical findings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.