acceptodds
Under review as a conference paper at ICLR 2027

On the Convergence of LoRA under Generalized Smoothness

Abstract

Low-rank adaptation (LoRA) has become a popularly used parameter-efficient approach for fine-tuning large pretrained models via replacing the full weight update with a factorized update . Despite its practical utility, existing convergence theory for LoRA remains limited: some prior analyses handle this factorized optimization problem by imposing boundedness assumptions on the LoRA factors, an additional assumption not guaranteed by LoRA updates. More recent results remove this assumption but still rely on -smoothness with respect to the weight variable . In practice, empirical studies of neural network training show that the Hessian norm can grow with the gradient norm, violating the -smoothness assumption on the bounded Hessian norm. Therefore, we analyze LoRA-GD and LoRA-SGD under generalized smoothness without assuming the bounded factor norms. For LoRA-GD, we prove an upper bound and the first lower bound, even when the stepsize is chosen adaptively from the optimization history. For LoRA-SGD with unbiased stochastic gradients of bounded variance, we obtain an high-probability guarantee. We also prove an lower bound at fixed confidence for adaptive stepsize policies that may depend on the observed optimization history and the current stochastic gradient. Experiments support our theoretical analysis: the Hessian norm grows with the gradient norm, the factor norms grow during optimization, and this growth amplifies stochastic factor-gradient noise.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.