acceptodds
Under review as a conference paper at ICLR 2027

Recoup the Experts: Coupling-Aware MoE Pruning under Deployed Routing

Abstract

Expert-pruning criteria for sparse Mixture-of-Experts (MoE) models are computed under the routing of the original model, but after pruning each token is routed to its top-k experts among the remaining ones and their gates are renormalized. We propose , a criterion that selects the set of experts whose outputs best reconstruct the layer outputs of the removed experts on the tokens they share. The objective has a closed form in the Gram matrix of gated expert outputs, and when expert outputs are orthogonal it reduces to a per-expert ranking that matches REAP up to the moment order. We also propose , which computes the same statistics under the routing of the pruned model. On four MoE models, the form that performs better at low budgets follows the off-diagonal mass of the Gram matrix, which can be computed before pruning. On Gemma-4-26B, the model with the highest off-diagonal mass, scores 59.2 on GSM8K when 25% of the experts are kept, against 15.6 for REAP. On Qwen3.6-35B, the model with the lowest, is 19.5 points above REAP on HumanEval+ when 12.5% are kept. These values are means over three disjoint calibration samples.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.