PoLoRA: A Preconditioned Orthogonalized LoRA Optimizer
Abstract
Low-rank adaptation (LoRA) reduces the cost of finetuning large language models by training low-rank updates instead of the full weight matrices. Each update is the product of two trainable matrix factors. These factors are usually trained with Adam, which effectively treats them as a single flat vector of parameters and ignores the matrix and product structure of LoRA. Applying a matrix-aware optimizer such as Muon to each factor does not consistently improve over Adam, and neither do the recently proposed product-aware variants of Muon. We introduce PoLoRA (Preconditioned Orthogonalized LoRA), a product-aware spectral optimizer that adds curvature preconditioning and controls the size of each factor update. We evaluate PoLoRA on instruction-tuning datasets for code and math across models from 1B to 8B parameters, and find that it reaches the final held-out loss achieved by tuned Adam with a – speedup in training steps and at most 3% per-step overhead. Compared to Adam, PoLoRA is also less sensitive to the learning rate, and its optimal learning rate is stable across ranks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.