acceptodds
Under review as a conference paper at ICLR 2027

AdamE: Dynamic Exponential Decay Rates for the Adam Optimizer

Abstract

Adaptive optimizers such as Adam are widely used in deep learning. Their gradient statistics are governed by two fixed decay rates, most commonly 0.9 and 0.999. The best pair differs across tasks, and searching for it is costly. We show that under any fixed second-moment rate the effective per-step gradient scale can shrink over training, the failure mode behind known counterexamples where Adam fails to converge. We address this with AdamE, a drop-in variant of Adam whose decay rates follow dynamic, closed-form schedules: the first rate starts at one half and decays smoothly toward zero, while the second rises toward one, with both rates determined by the iteration index. The schedules introduce no new hyperparameters and provably restore the monotonicity that fixed rates lack, recovering convergence without altering the update rule as AMSGrad does. We prove that AdamE attains near-optimal regret in the online convex setting, converges to a stationary point on smooth non-convex objectives, and admits an unconditional error bound on least squares, where the corresponding Adam guarantee holds only in a restricted regime. Experiments on language modeling, node classification, graph clustering, image classification, and image recognition show that AdamE achieves competitive convergence and final performance relative to well-tuned Adam-type baselines, with no manual tuning of the decay rates.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.