acceptodds
Under review as a conference paper at ICLR 2027

The Curvature Floor: A Geometric Criterion for Early Stopping

Abstract

Early stopping needs a clean validation set, which self-supervised pretraining and training with noisy labels often lack. Rules without one watch the loss, the gradient or the prediction changes, which keep changing after the model stops improving. We instead monitor the effective learning rate in units of the stability threshold of gradient descent, , where is the top eigenvalue of the Fisher information, preconditioned for adaptive methods and estimated with labels sampled from the model. In seven training regimes, from ResNet-50 on ImageNet to MAE pretraining, BERT fine-tuning and 40% label noise, falls by at least a factor of five and settles at a floor. It crosses 0.05 at 5% to 15% of the budget before the best checkpoint, a lead that varies almost 700-fold in updates, so a rule should wait for a time that scales with the budget. We prove that when the statistic depends only on the budget fraction, a threshold rule on a linearly smoothed statistic stops at the same fraction of every budget, for every profile of the statistic, if and only if its kernel is rescaled with the budget at every lag within the run. We also bound the delay of the exponential moving average in closed form. The resulting rule, Curvature Floor Stopping (CFS), uses no labels while it runs and adds about 5% memory on models up to 340M parameters. With four constants set per regime on development labels of the same runs, it matches or exceeds validation-based early stopping in five of seven regimes, stays within 0.02 points in the other two and cuts the total compute by 20% to 59%, monitoring included. With one budget-scaled time constant on held-out runs, it is on par with a stronger validation rule that returns its best checkpoint, while costing less than it in six of the seven regimes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.