acceptodds
Under review as a conference paper at ICLR 2027

Difficulty Aware Reweighting for Zeroth Order Fine-Tuning of LLMs

Abstract

Fine-tuning large language models (LLMs) is essential for downstream tasks, but backpropagation incurs prohibitive activation memory. Zeroth-order (ZO) optimization uses only forward passes, which provides a memory-efficient route to LLM adaptation on resource-constrained devices. Despite substantial progress on ZO optimizers and query schemes, the data-dependent characteristics of ZO gradient estimation remain under-explored. This work reveals that ZO gradient estimation exhibits sample-wise alignment heterogeneity w.r.t. the global descent direction, and hard samples yield more informative ZO updates for minimizing the global objective than easy ones. Based on this finding, we propose a difficulty-aware ZO (DaZO) algorithm that exploits sample-wise heterogeneity through sample reweighting. It down-weights easy samples while preserving informative hard samples. Further, we show that difficulty-aware reweighting reduces the global-gradient referenced estimation error and benefits faster convergence. Extensive experiments on RoBERTa-Large, Llama3-8B, and OPT models (1.3B/13B/30B) demonstrate significant improvement of DaZO over strong ZO baselines. Our work provides the first sample-centric view of ZO optimization, and shows that the quality of ZO gradient estimation depends systematically on the training samples. It is shown that difficulty-aware reweighting significantly improves ZO methods but has limited impact on FO methods, which highlights its particular relevance to the noisy gradient-estimation regime of ZO optimization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.