ParaMoLE: Efficient Tensor-Parallel Fine-Tuning of Mixture of LoRA Experts
Abstract
Mixtures of LoRA experts (MoLE) expand the adaptation capacity of large language models through sparse routing, but fine-tuning retains training states for the entire expert pool. Distributing these states can reduce memory use, yet memory feasibility alone does not ensure efficient training. Joint expert and router updates require cross-GPU coordination, while uneven expert batches can limit the efficiency of sparse computation. To address these challenges, we propose **ParaMoLE**, a framework for memory-efficient MoLE fine-tuning on tensor-parallel (TP) backbones under a fixed GPU budget. ParaMoLE adopts partial factor replication to reduce expert-state storage relative to full replication while jointly evaluating each selected expert within the backbone's TP group. During backward, ParaMoLE reuses reduced low-rank activations or activation gradients for both expert and router updates, avoiding separate synchronization of replicated-factor parameter gradients. Grouped matrix multiplication processes variable-sized expert batches in both forward and backward without padding them to a common length. We establish equivalence to the original MoLE outputs and gradients in exact arithmetic. Experiments across five NLP datasets show lower memory, training-step time, and communication on 3B and 14B LLMs, with comparable downstream quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.