SPARK: Sparse Spectral Adaptation Via Smooth Top- Selection
Abstract
Parameter-efficient fine-tuning (PEFT) methods such as LoRA, DoRA and, more recently, HiRA have become the standard way of adapting large language models to downstream tasks, as they train only a small fraction of the parameters and thereby make adaptation considerably faster and cheaper than full fine-tuning. However, every adapted task still produces an adapter that has to be stored and loaded at inference time, and in settings that require many adapters, for instance one per user, the accumulated storage and memory cost becomes a significant bottleneck. LoRA-XS addresses this problem by training only a small matrix placed between frozen factors obtained from the singular value decomposition (SVD) of the pretrained weight. Its adapters are extremely compact, but they are restricted to the leading singular directions, which are determined by the pretrained weight alone and are not necessarily the most relevant for the downstream task. In this paper, we propose SPARK, which removes this restriction. SPARK trains in a much larger subspace spanned by singular directions and simultaneously learns which of them to retain, using a differentiable top- selection that is gradually hardened during training. After training, only a matrix and the indices of the selected directions are stored, so the adapter is as small as that of LoRA-XS while being considerably more expressive. SPARK is compatible with any SVD-based update, and we apply it to both LoRA-XS and an SVD-based variant of HiRA. Experiments with Llama-3-8B on eight commonsense reasoning benchmarks show that SPARK consistently outperforms LoRA-XS-style adapters of the same size and surpasses HiRA while storing up to 88 times fewer parameters.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.