On Adaptivity in Zeroth-Order Optimization
Abstract
We investigate the effectiveness of adaptive zeroth-order (ZO) optimization for memory-constrained fine-tuning of large language models (LLMs). Contrary to prior claims, we show that adaptive ZO methods such as ZO-Adam provide no performance or convergence advantage over well-tuned ZO-SGD, while incurring significant memory overhead due to optimizer state. We provide theoretical justification for why Adam-like adaptivity is less effective in large-scale ZO settings, and confirm this through both synthetic and LLM fine-tuning experiments. Importantly, we show empirically that adaptivity can still provide a practical benefit in ZO optimization by improving robustness to step-size selection. To retain this robustness without incurring the memory cost of optimizer state, we propose MEAZO, a Memory-Efficient Adaptive ZO optimizer that tracks only a single scalar to adapt the global step size. We support our method with theoretical convergence guarantees under standard smoothness assumptions. Experiments on synthetic quadratic problems, multiple LLM families, and diverse fine-tuning tasks demonstrate that MEAZO performs comparably to competing methods while matching the robustness gains of ZO-Adam over ZO-SGD, without additional memory or latency overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.