AeroLoRA: You Need Few of Them
Abstract
Low-rank adaptation (LoRA) reduces the number of trainable parameters for fine-tuning large language models (LLMs), but two limitations remain: (i) maintaining separate adapters for multiple tasks incurs cumulative storage costs; and (ii) all low-rank channels remain active for every input, leaving input-dependent sparsity unexploited. We propose AeroLoRA (Adaptive Expert Routing with Orthogonal LoRA), a two-stage sparse adaptation framework that reduces both trainable parameter counts and the number of active adapter parameters per token. In the first stage, we derive a task-specific binary mask from the weight magnitudes of the trained up-projection matrix , restricting subsequent adaptation to the selected parameter subset. In the second stage, we update only the selected parameters and perform TopK routing based on activations from a frozen sparse projection, dynamically selecting a small subset of low-rank channels for each input without an additional trainable router. Experiments on Mistral and Llama across four benchmarks show that AeroLoRA achieves the best performance among the evaluated methods in single-task adaptation, multi-task adapter merging, and continual learning. During the second stage, trainable parameters account for only 0.24% and 0.22% of the total model parameters for Mistral and Llama, respectively, while AeroLoRA also activates the fewest adapter parameters per token among the evaluated methods. Spectral analysis shows that its adapter update matrices exhibit the highest effective rank among the compared methods. Subspace analysis further reveals that the update subspaces associated with TopK-selected channels align with the dominant directions of task gradients. These findings provide empirical evidence that diverse, task-relevant update directions and strong adaptation performance can coexist with very few trainable parameters and sparse, input-dependent channel activation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.