acceptodds
Under review as a conference paper at ICLR 2027

SLEP-MoEfication: Space-Level Expert Partition for Rich Information Preservation in Sparse Dense-to-MoE Conversion

Abstract

Sparse dense-to-MoE conversion methods cluster pre-trained dense feed-forward networks (FFNs) into selectively activated experts to avoid the high cost of pre-training Mixture-of-Experts (MoE) models from scratch. Calibration-based approaches guide expert construction using calibration dataset to maintain performance close to the pre-trained dense model at lower inference cost. However, consistent information preservation across the input datasets is limited by reliance on calibration-dependent representations. To address this limitation, we propose a space-level expert partition method for sparse dense-to-MoE conversion (SLEP-MoEfication). Specifically, we introduce a subspace coverage score that measures the fraction of dense weight energy captured by projection onto the subspace spanned by each expert's pre-trained neuron weights. We maximize the sum of subspace coverage scores to support information preservation while keeping pre-trained weight values fixed. Our method consistently outperforms the evaluated baselines across multiple models under a common training-free centroid-based router. Our empirical analyses reveal strong positive correlations between subspace coverage and information preservation, and between information preservation and task performance. These gains hold even when random partition outperforms other existing methods. Further analysis shows that our method improves the trade-off between expert activation sparsity and performance, reduces inference cost, and improves expert load balance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.