acceptodds
Under review as a conference paper at ICLR 2027

MSZO: Scheduled Multi-Scale Zeroth-Order Fine-Tuning of Large Language Models

Abstract

Compared with first-order (FO) backpropagation, zeroth-order (ZO) optimization is much more memory-efficient in fine-tuning large language models (LLMs) by estimating gradients using only forward passes. Existing ZO methods typically use a fixed perturbation scale for finite-difference gradient estimation throughout training, which implicitly assumes a stationary bias-variance trade-off across optimization stages. In this work, we show that this trade-off changes substantially during LLM fine-tuning. By probing checkpoints along LLM fine-tuning trajectories, we find a pronounced stage-dependent perturbation-scale effect: a small perturbation scale yields lower variance and lower FO-gradient referenced mean-squared error (MSE) early in training, but both can exceed those of a larger scale by the middle stage. In contrast, the larger scale provides more stable estimates at later stages at the cost of higher bias. Motivated by this finding, we propose MSZO, a scheduled multi-scale ZO optimizer that uses a small perturbation scale in the early stage and subsequently combines small- and large-scale estimates along the same random direction. This design dynamically balances bias and variance over the course of training. Experiments on RoBERTa-Large, OPT-1.3B, Llama3-8B, and Qwen3-4B demonstrate consistent improvement of MSZO over single-scale ZO fine-tuning. Unlike the common practice of using a fixed perturbation radius throughout fine-tuning, our results reveal that the appropriate scale in ZO optimization is inherently stage-dependent.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.