Understanding Learning-Rate Scaling with Batch Size in LoRA
Abstract
LoRA and full fine-tuning (FullFT) exhibit different learning-rate scaling behaviors as batch size increases, with LoRA requiring substantially larger learning rates at large batch sizes. The origin of this difference is not well understood and is often attributed to the low-rank nature of LoRA. We show that the gap reflects a learning-rate scaling mismatch rather than a capacity deficit. Starting from the coupled updates of the low-rank factors, we derive a state-dependent linear operator that characterizes the induced weight-space dynamics and the resulting effect on local curvature. This analysis reveals that LoRA changes the relative retention of gradient signal and stochastic noise, leading to a larger effective gradient-noise scale and consequently a different batch-size dependence of the optimal learning rate. Building on this mechanism, we derive a closed-form learning-rate scaling law that captures the dependence on batch size, rank, and model width. Fitted on the measured optima, the law accurately reproduces the observed scaling and, with no refitting, transfers to unseen models of different widths and architectures. A two-point calibration scheme derived from the law reduces learning-rate search cost by over 30.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.