acceptodds
Under review as a conference paper at ICLR 2027

Wanda-Guided Sparse Perturbations for Efficient Zeroth-Order Optimization

Abstract

Zeroth-order (ZO) optimization provides a memory-efficient alternative to backpropagation for fine-tuning large language models (LLMs), but its estimator variance grows with the ambient parameter dimension, often causing slow and unstable optimization. We propose Wanda-Importance Sparsity-MeZO (WIS-MeZO), a saliency-guided ZO method that concentrates sparse perturbations on task-relevant coordinates. WIS-MeZO uses weight magnitudes and downstream input-feature statistics to allocate the perturbation budget while keeping the pretrained model dense and requiring neither gradients nor pre-training data. We characterize support quality through the gradient energy capture ratio and show that, under stated conditions, WIS-guided supports preserve more gradient energy than random sparse supports. We further establish an -smooth convergence guarantee governed by the effective search dimension and captured gradient energy. Experiments across two model families, scales from 3B to 32B, and 16 benchmarks show improved aggregate performance over MeZO and random sparse perturbations, competitive or stronger results than magnitude-based Sparse-MeZO, and fewer optimization steps to reach comparable validation losses.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.