acceptodds
Under review as a conference paper at ICLR 2027

Zero-Order Optimization via Learnable Direction Sampling for LLM Fine-Tuning

Abstract

Zeroth-order (ZO) optimization estimates gradients through function evaluations along sampled perturbation directions. Standard isotropic sampling yields weak expected gradient alignment, limiting estimator quality and introducing dimension-dependent convergence rates. We propose a plug-and-play framework that learns the perturbation distribution to improve directional alignment. Our design follows from a descent bound that identifies expected gradient alignment as the objective for adapting the mean of a Gaussian sampling distribution. For an idealized directional algorithm with exact policy updates, we characterize the evolution of alignment and derive convergence guarantees with no explicit dimension dependence. Building on this formulation, we develop a practical finite-difference algorithm that updates the sampling policy using REINFORCE estimates and integrates with existing ZO optimizers. The resulting framework is particularly relevant to large language model (LLM) fine-tuning, where avoiding backpropagation reduces training memory requirements. Experiments demonstrate consistent improvements across models, ZO optimizers, and fine-tuning schemes under matched forward-evaluation budgets, supporting learned direction sampling as an effective mechanism for improving ZO fine-tuning. Our code is available at [https://anonymous.4open.science/r/zo_ldsd](https://anonymous.4open.science/r/zo_ldsd).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.