acceptodds
Under review as a conference paper at ICLR 2027

HiZO: Hebbian-informed Zeroth-order Optimization for Efficient Fine-tuning

Abstract

Zeroth-order optimization enables language model fine-tuning without backpropagation, but isotropic perturbations do not exploit the task-dependent information contained in forward activations. We introduce Hebbian-informed Zeroth-order Optimization (HiZO), which uses this information to guide the perturbation distribution and improve convergence. HiZO uses the magnitude of centered input–output covariance to allocate perturbation variance across weight coordinates while preserving the total variance budget. The resulting guidance reflects both learned weights and downstream input statistics. These statistics are collected during existing loss evaluations and require no additional model forward passes. We also develop subsampled and separable variants to reduce the computation and storage needed for guidance. For smooth nonconvex objectives, we establish convergence bounds in the directional-derivative limit that characterize the roles of covariance–gradient alignment and estimator variability. A complementary one-layer analysis identifies conditions under which Hebbian covariance provides information about task-gradient magnitudes. Language-model experiments and controlled ablations show faster optimization with Hebbian guidance while retaining comparable or better predictive performance. These findings suggest that forward activations offer a practical source of task-dependent information for improving zeroth-order convergence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.