acceptodds
Under review as a conference paper at ICLR 2027

Understanding Clipping in Zeroth-order Optimization

Abstract

Zeroth-order optimization has emerged as a resource-efficient alternative to Stochastic Gradient Descent (SGD) when fine-tuning LLMs, thanks to the inherent low dimensionality of the loss landscape. It relies on a two-point estimate of the gradient that only accesses the function value and not the gradient, bypassing the resource-heavy backpropagation. Recently, it was accidentally discovered that per-sample clipping can improve the performance of the trained model, while running a differentially private version of zeroth-order optimization zhang2024dpzero with a very large privacy parameter (corresponding to very little privacy). In this paper, we systematically investigate this phenomenon and demonstrate that () the optimal choice of learning rate is inversely proportional to the clipping threshold , satisfying constant; () clipping drives the zeroth-order optimization to a different solution with higher test accuracy; and () this phenomenon is only observed with per-sample clipped zeroth-order optimization and not other zeroth-order alternatives that also provide robust estimates of the gradient. Further analyses with artificially injected label noise reveals that clipping leads to more robust predictions, which stems from the implicit down weighting of hard-to-fit training examples.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.