Dynamic Tuning of Weight Decay: Stability Analysis and Empirical Scheduling
Abstract
Weight decay is often held constant while the learning rate varies during training, although both affect optimization dynamics. In this paper, we investigate this interaction on two levels: static hyperparameter coupling between training configurations, and the potential performance gains from dynamic weight decay scheduling. Specifically, under uniform stability analysis, we characterize how weight decay contracts perturbations in stochastic gradient descent (SGD) and changes the learning-rate conditions for stability. Based on this, matching stability bounds under a heuristic regularization-equivalence hypothesis suggests the coupling , relating weight decay , learning rate , and the number of updates . For a fixed number of training epochs, this coupling further suggests that the product scales with batch size . Motivated by this inverse relationship, we examine two weight-decay schedules, linear-up and iso, that increase as decays while leaving the learning-rate schedule unchanged. In our experiments, hyperparameter sweeps on image-classification tasks and in Qwen3-0.6B for LoRA fine-tuning show trends consistent with the proposed coupling. Scheduling experiments across architectures and tasks with SGD show higher reported accuracy with increasing weight-decay schedules than with fixed weight decay in several settings. These findings support considering weight decay jointly with the learning rate when choosing its magnitude and scheduling it during training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.