Adaptive Zeroth-Order Optimization via Sample-Wise Second Moments
Abstract
Zeroth-order optimization (ZO) estimates gradients using function evaluations without requiring access to first-order derivatives. Existing ZO methods either rely on isotropic perturbations or require additional function evaluations to obtain adaptive curvature information, introducing extra computation overhead and slowing the convergence. We propose two adaptive ZO optimizers that share an efficient diagonal second-moment estimator built from per-sample gradient surrogates. This state shapes both perturbations and updates, with distinct update scalings and shared adaptive damping for stability and perturbation-scale control. We provide convergence proofs for both variants and demonstrate the first variant’s superior performance on test functions with heterogeneous curvature. Experiments on both full-model and parameter-efficient language-model fine-tuning show that standard and adaptive ZO baselines require up to as many forward passes as our methods, while our methods achieve competitive final performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.