Improving Low-Rank Adaptation by Adjusting its Dynamics
Abstract
Low-Rank Adaptation (LoRA) efficiently fine-tunes pre-trained large language models with two trainable low-rank matrices B and A. Previous LoRA variants primarily improve performance by designing the initialization or update dynamics of and to preserve more information from the full parameter gradient matrix. Such information preservation depends on the alignment between the subspaces induced by and the current gradient, which is difficult to guarantee throughout training. Nevertheless, few works investigate the conditions on and that yield the most loss descent for a given current adapted weight at an arbitrary training iterate. In this paper, we theoretically analyze the loss descent of LoRA from the -th step to the -th step and provide the conditions on and for the most loss descent: the largest and smallest singular values of and equal the square roots of the singular values of with aligned internal singular subspaces. We reveal that the balance is a simple and easy-to-implement case of the conditions on and , and propose the method Balance-Preserving Dynamics Adjustment for LoRA (BDA-LoRA) by adjusting the dynamics to make and balance in each step. Specifically, BDA-LoRA obtains a gradient-approximate balanced initialization via LoRA-based randomized SVD without computing or storing the full parameter gradient matrix, and modifies raw optimizer update directions slightly to make and balance from and for all . Experiments on NLU, NLG and commonsense reasoning tasks show the effectiveness and superiority of BDA-LoRA.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.