Label Smoothing Improves Gradient Ascent in LLM Unlearning
Abstract
LLM unlearning has emerged as a promising approach, aiming to enable models to forget hazardous/undesired knowledge at low cost while preserving as much model utility as possible. Among existing techniques, the most straightforward method is performing Gradient Ascent (GA) w.r.t. the forget data, thereby forcing the model to unlearn the forget dataset. However, GA suffers from severe instability, as it drives updates in a divergent direction, often resulting in drastically degraded model utility. To address this issue, we propose Smoothed Gradient Ascent (SGA). Theoretically, we analyze how the smoothing rate affects the update direction and derive a local condition for reducing its component along the pure GA direction. Intuitively, this extends GA by combining gradient contributions from forget and normal data, with the smoothing rate controlling their relative weights. Empirically, we evaluate SGA on three benchmarks: TOFU, Harry Potter, and MUSE-NEWS. Experimental results show that, with an appropriate smoothing rate, SGA can improve the forgetting–utility trade-off over GA and achieve competitive performance against existing baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.