MAES: Modality-Aware Expert Slimming for Multimodal Mixture-of-Experts
Abstract
The rapid adoption of Mixture-of-Experts (MoE) in multimodal large language models has enabled remarkable capabilities, yet at the cost of critical bottlenecks in storage, memory, and inference speed. Structural pruning offers fine-grained compression that permanently reduces memory and computation, but existing methods treat all tokens uniformly and ignore the fundamental heterogeneity between modalities. Through systematic analysis, we uncover two key phenomena in multimodal MoEs: (i) experts exhibit pronounced modality affinity, inducing modality-specific redundancy; and (ii) different modalities exhibit different activation magnitude distributions, rendering modality-agnostic channel scoring unreliable. These findings motivate our Modality-Aware Expert Slimming (MAES), a training-free structural pruning and inference framework that proposes Expert Modality Affinity (EMA)-guided budget allocation according to each modality's contribution. To rank expert redundancy, we further introduce an exact second-order expert-importance criterion, replacing noisy proxies used in prior work. For deployment, we rearrange experts jointly across layers to balance weight-memory load across GPUs, which seamlessly integrates with modern inference frameworks and preserves efficient expert parallelism. Across 12 image and video benchmarks and 3 model families, MAES consistently outperforms existing MoE pruning methods, retaining over 98% of original performance at 30% channel pruning while attaining lower memory usage and higher throughput, achieving the state-of-the-art accuracy–efficiency trade-off. Our code is released at https://anonymous.4open.science/r/MAES.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.