LoSARE: Low-Rank Optimization for Training LLMs with Adaptive Subspace Tracking
Abstract
Low-rank optimization has emerged as an effective approach to reducing the memory cost for training Large Language Models (LLMs). Existing methods are still insufficient in adapting to evolving gradient subspaces and recovering gradient information discarded by low-rank projection. We present LoSARE, a memory-efficient low-rank optimization method that addresses both issues through adaptive subspace tracking and energy-aware residual enhancement. LoSARE incrementally updates the projection subspace according to the current gradient geometry. The gradient subspace is continually tracked with low-complexity QR decomposition, avoiding the need for costly SVD re-computation in GaLore. It further enhances the discarded gradient residual according to its relative energy, while low-rank optimizer states and stable norm-growth control are still retained. Extensive experiments on LLaMA models demonstrate that LoSARE consistently outperforms GaLore and Fira in memory-efficient pre-training and fine-tuning, under the constraints of preserving their low memory footprint. Notably, on the LLaMA-7B pre-training task, LoSARE with projection rank 64 achieves comparable or better performance than GaLore with ranks 512 and 1024, using and smaller ranks, respectively. At the same time, LoSARE achieves and faster training speed than GaLore with rank 1024 and Fira with rank 64, respectively. Furthermore, LoSARE maintains strong performance under substantially lower rank constraints and exhibits more effective late-stage optimization. These results demonstrate that our method provides an effective balance among performance, memory efficiency, and training cost for LLM training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.