acceptodds
Under review as a conference paper at ICLR 2027

ConMoE: MoE compression as expert selection and functional assignment

Abstract

Sparse Mixture-of-Experts (MoE) models scale model capacity through sparse activation, but storing their full parameters remains costly in GPU memory. Existing expert compression methods address this problem through pruning, merging, or substitution, typically coupling expert pool construction with the assignment of routed workloads. In this work, we formulate expert compression as two decisions—constructing an expert pool and assigning the original experts’ workloads to that pool—and show that assignment can be optimized after the pool is fixed. Based on this formulation, we introduce **ConMoE**, which constructs a compact pool using Anchor Coverage and optimizes assignment through a router-conditioned functional objective derived from a layer reconstruction bound, treating parameter proximity only as a proxy. ConMoE uses unlabeled calibration inputs and keeps the router and expert weights frozen. With the pool fixed, functional assignment improves macro accuracy by 1.82 and 0.35 percentage points at 50% and 75% expert retention, respectively, and also improves pools constructed by other selectors. Across three MoE architectures and six zero-shot benchmarks, ConMoE achieves the highest macro average among the compared methods at both ratios. At 50% expert retention, ConMoE reduces peak GPU memory by 29.6–48.4% relative to the uncompressed models on our generation workloads. These results establish expert assignment as an independent optimization axis for MoE compression.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.