LLM2llm: Parameter-Efficient Adaptation via Transferred Experts from MoE LLMs
Abstract
Stronger language models acquire rich knowledge through expensive pre-training, but transferring it typically requires learning from teacher-generated supervision. We propose LLM2llm, a framework that extracts reusable expert components from a stronger source Mixture-of-Experts model for direct inheritance by target pretrained model. LLM2llm identifies cross-domain generalist experts using multi-domain routing statistics and condenses them through low-rank subspace extraction and score-weighted fusion. This one-time extraction requires no additional training, and the resulting experts can be reused across multiple downstream tasks. Injected as residual branches, only the transferred experts are updated during downstream supervised fine-tuning; the original target parameters remain frozen, and the source model is no longer required. Across two heterogeneous source–target settings, LLM2llm achieves higher average performance than Full Fine-Tuning and representative parameter-efficient baselines on language understanding and dialogue generation tasks. Ablations support the effectiveness of expert selection, condensation, and transferred initialization. These results demonstrate the potential of directly inheriting reusable model components to improve downstream learning efficiently.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.