Embedding Noise Should Follow the Learning Rate in LLM Fine-Tuning
Abstract
Adding random noise to token embeddings improves supervised fine-tuning (SFT) of large language models (LLMs). The noise helps early in training, when a small instruction set is revisited for several epochs, but it becomes a problem at the end. As the learning rate (LR) decays, the optimizer takes smaller steps, but noise of fixed scale keeps adding a gradient bias that does not shrink. In our case study, keeping the noise through the last quarter of training leaves the model farther from a stationary point of the clean training loss, which is the loss without noise, and the gradient of this loss stays larger than when the noise is removed. We introduce AdaNoiseSFT, which sets the variance of the embedding noise proportional to the current LR, so the noise is strong while the LR is high and fades as the LR decays, with no additional model passes or optimizer state. Our analysis for Adam assumes that the smoothing effect of the noise does not vanish late in training. Under this condition, fixed noise leaves a gradient bias that does not vanish for any LR schedule, whereas coupling the noise to the LR drives this bias to zero when the LR decays to zero. To test that the benefit comes from following the LR schedule itself, we evaluate cosine, linear, step, and warmup stable decay (WSD) LR schedules against independently tuned noise tapers that follow elapsed training time. On Llama-3.1-8B with Alpaca-cleaned and a cosine LR, AdaNoiseSFT reaches 19.5 length-controlled AlpacaEval 2.0 (AE-2) over five seeds, versus 18.1 for the strongest baseline, a time taper selected on a separate selection split. Over ten seeds under linear, step, and WSD LR schedules, it beats the selected time taper by 1.39, 1.14, and 0.91 AE-2 points, respectively. For all four LR schedules, the best noise schedule is the one that follows the shape of the LR.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.