EAD-LoRA: Expert-Aware Differential Initialization for MoE-LoRA Frameworks
Abstract
Parameter-efficient fine-tuning (PEFT), particularly Low-Rank Adaptation (LoRA), has become essential for adapting large foundation models to downstream tasks. To enhance representation capabilities, Mixture-of-Experts (MoE) architectures are often combined with LoRA (MoE-LoRA) by employing multiple low-rank adapters. However, existing MoE-LoRA methods uniformly initialize all experts using random distributions, ignoring the intrinsic structure of pretrained models and the anticipated functional divergence among experts. We propose EAD-LoRA, Expert-Aware Differential Initialization for MoE-LoRA frameworks. Unlike uniform initialization, our method uses a brief warm-up phase to collect routing data, then applies Principal Component Analysis (PCA) to initialize each expert's parameters based on its specific input distribution. Extensive experiments demonstrate that EAD-LoRA converges faster and achieves an average accuracy improvement of 1.6% and 5.8% over standard MoE-LoRA on different backbones, while its hybrid variant EAD-LoRA+ brings up to 2.7% and 7.7% improvements.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.