Adaptive Stepsizes Meet Momentum: An Improved Zeroth-Order Method
Abstract
Training deep networks relies heavily on adaptive step sizes and momentum mechanisms. Yet, in gradient-free regimes, their theoretical foundations are still underdeveloped. Since many problems permit only zeroth-order oracle access, classical gradient-based analyses cannot be applied. To address this, we propose ZO-AdaNesterov, a stochastic zeroth-order Nesterov momentum method equipped with adaptive step sizes, and develop a rigorous convergence analysis. We derive oracle complexity bounds for both convex and nonconvex objectives, which are strictly sharper than those of existing zeroth-order adaptive momentum methods. Besides, we provide the *first* last-iterate saddle-avoidance result for zeroth-order adaptive methods: with probability one, ZO-AdaNesterov does not converge to strict saddle points. Numerical experiments further confirm the efficiency of our method.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.