On the Convergence of Stochastic Low-Rank Adaptation
Abstract
Low-rank adaptation (LoRA) minimizes the training loss after adding a trainable low-rank adapters to a frozen pretrained weight matrix. Previous analysis of gradient descent for LoRA (LoRA-GD) gives a sufficient complexity of iterations to find an -stationary point in the deterministic setting. We sharpen this analysis and show that iterations suffice for the same first-order criterion. We further study stochastic LoRA under unbiased stochastic gradients with finite variance. We propose LoRA-NSGDM, which finds an -stationary point with stochastic oracle calls. Under an additional mean-square smoothness condition, we propose LoRA-STORM, which improves the stochastic oracle complexity to .
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.