No Expert Replaces Another: Pruned Experts' Outputs Are Not Recovered by the Remaining Ones
Abstract
Expert pruning, merging and replacement in sparse Mixture-of-Experts (MoE) models are commonly motivated by expert redundancy, inferred from output similarity on shared inputs, from the small loss change after zeroing a single expert, or from replacing an expert by its mean output. None of these measurements tests whether the remaining experts can reproduce a pruned expert's output on the tokens routed to it, which is where pruning changes the layer output. We measure this on four MoE models. The remaining experts, combined linearly, recover only 2 to 28% of the pruned output energy, whereas a noisy copy of the pruned expert recovers 98%. Pruning still works at moderate budgets because this energy is concentrated in few experts: keeping half of them removes only 4.5 to 18% of it, and within a budget the energy a criterion removes predicts its perplexity. Three consequences follow. Pruning scores should be summed over an expert's routed tokens rather than averaged, which lowers perplexity in 15 of 18 settings on the three larger models. Modeling interactions between experts changes the selection little. Calibration data should match the deployment domain rather than be large: 2K matched tokens outperform two million mismatched ones.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.