GaNsZO: Gauss-Newton Step-sized Zeroth-Order Optimization for Memory-Efficient LLM Fine-Tuning
Abstract
Fine-tuning large language models at scale is increasingly constrained by the memory cost of backpropagation. Zeroth-order (ZO) optimizers (e.g., MeZO) alleviate this bottleneck by relying solely on forward passes, but they often suffer from slow convergence. Recent works such as HiZOO utilize diagonal Hessian estimators as preconditioner for acceleration, yet this incurs additional memory costs comparable to training parameters. Moreover, the noisy diagonal Hessian estimates often lead to training instabilities. In this work, we introduce the uss-ewton tepsized eroth-rder (GaNsZO) algorithm, a ZO optimizer. GaNsZO estimates the directional second derivative along the descent direction and uses it as a stepsize scaling factor, reducing memory overhead while improving robustness. To estimate efficiently, GaNsZO exploits the model's composite structure and computes Gauss-Newton quadratic form along that direction, saving one forward pass compared to standard central-difference schemes. Empirically, on OPT and Llama models, GaNsZO consistently outperforms HiZOO in terms of accuracy and training stability. In certain tasks (e.g., SST-2), GaNsZO approaches the performance of first-order fine-tuning. In terms of efficiency, GaNsZO reduces memory consumption by up to 43% compared to HiZOO, while requiring significantly fewer forward passes to reach comparable loss, leading to up to end-to-end speedup in selected settings. Finally, we show that GaNsZO combines naturally with PEFT and recent ZO techniques (e.g., LOZO) and produces further improvements.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.