acceptodds
Under review as a conference paper at ICLR 2027

Router-Centric Adaptation for MoE LLMs

Abstract

Existing parameter-efficient fine-tuning (PEFT) methods for Mixture-of-Experts (MoE) LLMs primarily introduce trainable parameters into experts, yet ignore the router. In this paper, we find that, under top- expert selection, models adapted with these methods largely inherit their token-to-expert assignment patterns from pre-training. Based on our empirical observations, we argue that adapting routing preferences is crucial for downstream adaptation. We therefore propose Router-Centric Adaptation (RoCA), a PEFT method that accounts for the routing characteristics of MoE LLMs. Its core strategy is to explore potentially task-relevant experts to help shape task-adaptive routing preferences and improve expert utilization. It further uses routing uncertainty to adaptively inject domain knowledge. Experiments across three recent MoE backbones (Qwen3-30B-A3B, DeepSeek-V2-Lite, and OLMoE-1B-7B) and five domains show that RoCA achieves the highest average downstream performance while using 5–10× fewer trainable parameters than the second-best method, demonstrating its efficiency and effectiveness. RoCA also improves general capabilities after fine-tuning. We hope our work provides a new perspective on PEFT for MoE LLMs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.