acceptodds
Under review as a conference paper at ICLR 2027

GS-MoE: Training-Free MoE Compression via Groupwise Subspace Sharing and Pruning

Abstract

Mixture-of-Experts (MoE) models reduce computation through sparse activation, but storing their expert parameters remains a major memory bottleneck. We introduce GS-MoE, a training-free compression framework that addresses redundancy across and within experts through two compression axes: shared-subspace rank and expert intermediate width. Groupwise low-rank representations capture shared structure across experts while retaining expert-specific projections, and Nystr\"om-based neuron compression reduces intermediate width within each expert. For fixed shared subspaces, we derive per-expert retained-energy gains from the reconstruction objective and use these gains to guide expert grouping. A gradient-free layer-importance score guides the allocation of rank and width budgets across layers, with non-uniform ranks across groups and intermediate widths across experts. Experiments show that compression along both axes achieves higher accuracy than compression along either axis alone. On DeepSeek-MoE-16B-Base, at a 60% reduction in routed-expert parameters, GS-MoE retains over 92% of the uncompressed model's average reasoning performance, compared with 73% for D-MoE, while achieving 30% lower out-of-domain perplexity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.