Zero-Order Optimization via Learnable Direction Sampling for LLM Fine-Tuning
Abstract
Zeroth-order (ZO) optimization estimates gradients through function evaluations along sampled perturbation directions. Standard isotropic sampling yields weak expected gradient alignment, limiting estimator quality and introducing dimension-dependent convergence rates. We propose a plug-and-play framework that learns the perturbation distribution to improve directional alignment. Our design follows from a descent bound that identifies expected gradient alignment as the objective for adapting the mean of a Gaussian sampling distribution. For an idealized directional algorithm with exact policy updates, we characterize the evolution of alignment and derive convergence guarantees with no explicit dimension dependence. Building on this formulation, we develop a practical finite-difference algorithm that updates the sampling policy using REINFORCE estimates and integrates with existing ZO optimizers. The resulting framework is particularly relevant to large language model (LLM) fine-tuning, where avoiding backpropagation reduces training memory requirements. Experiments demonstrate consistent improvements across models, ZO optimizers, and fine-tuning schemes under matched forward-evaluation budgets, supporting learned direction sampling as an effective mechanism for improving ZO fine-tuning. Our code is available at [https://anonymous.4open.science/r/zo_ldsd](https://anonymous.4open.science/r/zo_ldsd).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.