LoLoRA: Locally Fine-Tuned Low Rank Adapters
Abstract
Low-Rank Adaptation (LoRA) is a memory-efficient fine-tuning method for large language models (LLMs) that approximates weight updates as , where , and . To maximize memory savings, one can freeze matrix , avoiding the storage of its input activations, but this often degrades performance. In this work, we mitigate this trade-off by introducing gradient-free updates to matrix during the forward pass. Our method computes these updates based on the layer’s immediate input, allowing it to adapt to input distribution shifts without storing activations for the backward pass. Experiments show that our approach achieves fine-tuning accuracy comparable to standard LoRA, while reducing peak memory consumption by up to 20% on LlaMA3-8B.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.