acceptodds
Under review as a conference paper at ICLR 2027

How Smoothness Controls Overfitting: Sharp Curvature–Time Laws for Gradient Descent

Abstract

How does the population risk of gradient descent evolve as optimization continues on the same finite training sample? Classical stability bounds can overstate the statistical cost of longer training. We establish sharp curvature–time laws for projected gradient descent on the unit ball with convex, -Lipschitz, -smooth losses for every . At unit smoothness, the worst-case expected excess risk is for sample size and cumulative step size . Although the classical stability certificate is already constant-order at , the worst-case risk remains at the optimal scale and reaches constant order only at . The key is an interaction–dissipation principle that balances higher-order sample interactions against the stabilizing effect of losses unchanged by sample replacement. Beyond convexity, sharp risk laws show that even vanishing negative curvature can separate the last iterate from the full parameter average. On a single weakly nonconvex instance, the last iterate, a uniformly sampled iterate, and the full parameter average all have risk at an earlier horizon, but at a common later horizon their risks are , , and , respectively. Further applications yield sharp guarantees for weakly regularized pairwise learning, randomized coordinate descent, and fixed-source learning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.