UltraSparse: Extreme Compression of Upcycled MoE Models
Abstract
Mixture-of-Experts (MoE) has become a promising architecture for scaling large models through sparse activation. Recently, upcycling dense models into MoE models has emerged as an effective approach for constructing powerful sparse models, motivating the need for effective compression techniques for upcycled MoE models. Recent studies reveal that upcycled MoE models exhibit an expert-shared base weights and expert-specific delta weights structure, where the delta weights contain substantial redundancy and compression potential. However, how to achieve extreme compression of expert-specific delta weights while preserving model capability remains largely unexplored. In this work, we conduct the first systematic study on the extreme sparsification of expert-specific delta weights and reveal that both weight distribution and knowledge distribution play critical roles in governing model performance under ultra-high sparsity regimes. Building upon these insights, we present UltraSparse, the first framework for extreme compression of upcycled MoE models. A Weight Distribution Preservation module is proposed to maintain consistency between the weight distributions before and after sparsification. A Knowledge Distribution-Guided Hybrid Sparsification module is proposed to preserve critical knowledge associated with high-magnitude weights while retaining complementary knowledge contributed by low-magnitude weights. To comprehensively evaluate compression methods for upcycled MoE models, we establish a large-scale benchmark covering diverse upcycling approaches, model architectures of different scales, and datasets spanning different tasks and modalities. Extensive experiments demonstrate that UltraSparse consistently outperforms existing compression methods across various settings, and achieves unprecedented compression ratios, reaching up to 1000× compression while maintaining strong model performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.