acceptodds
Under review as a conference paper at ICLR 2027

AdaUpMoE: Adaptive Fine-grained Compression of Upcycled MoE Models

Abstract

Mixture-of-Experts (MoE) has become a promising architecture for scaling large models with sparse activation. Recently, upcycling dense models into MoE models has emerged as an effective approach for constructing powerful sparse models, motivating the need for effective compression techniques for upcycled MoE models. In this work, we present the first systematic study of upcycled MoE compression and reveal two key observations: (1) existing methods typically apply uniform compression strategies across the entire model, overlooking the heterogeneous compression requirements among different MoE components; (2) compression effectiveness is highly sensitive to fine-grained compression policies, including sparsity ratios and rescaling factors. Based on these findings, we propose AdaUpMoE, an adaptive fine-grained compression framework for upcycled MoE models. Specifically, AdaUpMoE formulates compression as a component-wise policy learning problem, leveraging a small amount of unlabeled data and a two-stage distillation-based learning paradigm to automatically learn compression policies for different layers, experts, and model modules. By adaptively allocating compression budgets according to component characteristics, AdaUpMoE effectively exploits the heterogeneous redundancy of upcycled MoE models while preserving model capability. Extensive experiments on 5 upcycled MoE approaches, 8 models with diverse architectures and scales, and 30 datasets covering various tasks and modalities demonstrate that AdaUpMoE consistently outperforms existing compression approaches, achieving particularly significant improvements under extreme (up to 100× on expert-specific deltas) compression regimes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.