ConMoE: MoE compression as expert selection and functional assignment
Abstract
Sparse Mixture-of-Experts (MoE) models scale model capacity through sparse activation, but storing their full parameters remains costly in GPU memory. Existing expert compression methods address this problem through pruning, merging, or substitution, typically coupling expert pool construction with the assignment of routed workloads. In this work, we formulate expert compression as two decisions—constructing an expert pool and assigning the original experts’ workloads to that pool—and show that assignment can be optimized after the pool is fixed. Based on this formulation, we introduce **ConMoE**, which constructs a compact pool using Anchor Coverage and optimizes assignment through a router-conditioned functional objective derived from a layer reconstruction bound, treating parameter proximity only as a proxy. ConMoE uses unlabeled calibration inputs and keeps the router and expert weights frozen. With the pool fixed, functional assignment improves macro accuracy by 1.82 and 0.35 percentage points at 50% and 75% expert retention, respectively, and also improves pools constructed by other selectors. Across three MoE architectures and six zero-shot benchmarks, ConMoE achieves the highest macro average among the compared methods at both ratios. At 50% expert retention, ConMoE reduces peak GPU memory by 29.6–48.4% relative to the uncompressed models on our generation workloads. These results establish expert assignment as an independent optimization axis for MoE compression.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.