acceptodds
Under review as a conference paper at ICLR 2027

Reconciling Routing and Learning for Mixture-of-Experts Expansion

Abstract

Expert expansion adds task-specific capacity to mixture-of-experts (MoE) models without increasing inference-time computation. However, experts receive task gradients only from tokens assigned by pretrained routing, which may exclude critical downstream learning signals and ultimately limit task performance. In this paper, we propose ReconExp, a proactive expansion approach that addresses this misalignment in both expert initialization and training. Specifically, ReconExp uses downstream gradients to rank pretrained experts and initialize the new expert from the highest-ranked expert. ReconExp also distills pretrained expert outputs to expand learning beyond the inherited routing region. Together, these designs make the new expert responsive to downstream learning rather than solely to pretrained routing. Experiments show that ReconExp improves the accuracy on BFCL over the strongest expansion by 5.3 points and reaches 97% of the Full-Finetune score on WebShop, while updating only 1.45% of the model parameters.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.