ZOLD: A Zeroth-Order Langevin Dynamics Framework for LLM Fine-Tuning
Abstract
Zeroth-order (ZO) optimization has emerged as a memory-efficient paradigm for fine-tuning LLMs, estimating gradients from forward evaluations alone. However, existing approaches refine the ZO gradient estimator within an SGD-style update and thereby inherit a structural limitation of SGD: the absence of a mechanism for global exploration of the loss landscape. In this paper, we propose a novel framework, Zeroth-Order Langevin Dynamics (ZOLD), which can be viewed as a zeroth-order analogue of stochastic gradient Langevin dynamics. A precise coupling between the smoothing radius and the inverse temperature ties the ZOLD dynamics to a Gibbs measure concentrating on the global flat minima of the loss, i.e., the global minimizers of a Hessian-trace regularized objective. Within this framework, we further introduce layerwise normalized ZOLD (LN-ZOLD) that addresses the layerwise scale heterogeneity of pretrained LLMs at no additional memory or forward evaluation. On the theoretical side, we establish non-asymptotic Wasserstein-1 error bounds for ZOLD with respect to the flatness-biased Gibbs measure. Empirically, ZOLD and LN-ZOLD improve over MeZO in accuracy across OPT-13B/30B/66B, LLaMA3, and Phi-2, while remaining substantially faster than HiZOO and FZOO at comparable accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.