Multiplicative Mixture of Experts for Rank-Efficient LLM Finetuning
Abstract
Large language models (LLMs) have achieved impressive results in many general-purpose domains, but their performance on specific tasks can still be improved through finetuning. Parameter-efficient finetuning (PEFT) tailors an LLM to one or more tasks through a small amount of trainable parameters, requiring reduced computational resources. On one hand, techniques like low-rank adaptation (LoRA) provide the required parameter efficiency with additive adapters of low, and fixed, rank, which limits their flexibility. On the other hand, mixture of experts (MoEs) enhance the flexibility of a model at the cost of parameter count and memory budget. The combination of the two approaches, parameter-efficient MoEfication, has shown promise in addressing the issues of both. In this work, we show that {replacing additive with multiplicative interactions improves the efficiency of PEFT adapters, increasing the flexibility and reducing the number of parameters involved in MoEfication. We find that our quantum-inspired multiplicative method, OperA, is optimal given the same parameter budget for 7/8 models considered, using fewer or equal parameters than the baseline. In a multi-task setting, OperA fully surpasses all baselines on all models. We also propose a class of DMRG-based algorithms to accelerate OperA in large-scale deployments by alleviating its computational complexity, resulting in up to a 2.5x throughput improvement and up to a 8.9x memory reduction over the naive implementation. Finally, we provide evidence that OperA surpasses the rank of competing solutions by more than two orders of magnitude.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.