LoREAL: Low-Rank Expert-Input Adaptation for Mixture-of-Experts Language Models
Abstract
Mixture-of-Experts (MoE) language models use the same token representation to select experts and to compute their outputs. Downstream adaptation can therefore act on two distinct pathways: routing and expert computation. We propose Low- Rank Expert Adaptation (LREA), which separates these pathways and adapts only the inputs delivered to the experts. At each MoE layer, the frozen routing pathway uses the layer input to select experts and compute mixture weights. In parallel, the adaptive expert-input pathway passes the same input through a shared low-rank residual adapter before sending it to the selected frozen experts. The local adapter thus changes what the experts receive without changing the input used for routing at that layer; earlier adapted layers may still affect later routing decisions. The shared adapter introduces 2dr trainable parameters per layer, with optional per-expert scaling adding N scalars. On OLMoE-1B-7B, LREA outperforms router tuning across continued pretraining and supervised fine-tuning, reduces measured forgetting relative to router tuning, and improves on expert LoRA in most comparisons at comparable parameter budgets. Ablations support adapting the expert-input pathway as an effective alternative to increasing router capacity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.