GRAFT-MOE: GROUP-GUIDED ROUTING ADAPTA- TION AND FUNCTIONAL TRANSFER FOR PRETRAINED MOE EXPANSION
Abstract
Mixture-of-Experts (MoE) expansion must add capacity without disrupting pretrained routing. We propose GRAFT-MoE, which groups pretrained experts using routed-input statistics and shared-probe responses, allocates new capacity by group traffic, selects representative parents, and applies a layerwise routed-output warm-start that jointly updates new experts and the router. During continued training, confidence-aware intra-cluster exploration allows low-confidence experts to be replaced by same-group alternatives while preserving distinct selections; it is disabled at inference. On OLMoE-1B-7B-0125 and Qwen1.5-MoE-A2.7B at , , and expansion, GRAFT-MoE achieves the best average downstream score in all six settings and the lowest perplexity in five. At , it reaches 8.86 PPL / 65.08 Avg on OLMoE and 10.97 PPL / 66.63 Avg on Qwen, and recovers the source NLL after 8.39M tokens (12.17 GPU-hours) versus 12.58M (16.29 GPU-hours) for EU-GN. It improves over EU-GN by 0.29/0.36 average-score points and over full-parameter continued pretraining by 0.79/0.77 points on OLMoE/Qwen at expansion, while reducing recovery cost by 25.3%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.