MAP-MoE: Modality-Aware Structured Channel Pruning for Multimodal Mixture-of-Experts Models
Abstract
Sparse Mixture-of-Experts (MoE) models improve performance by scaling up while activating only a few experts for each token. This scaling advantage requires all experts to be stored during deployment, leading to substantial storage and GPU memory costs. Existing solutions reduce these costs by pruning experts, merging experts, or pruning expert feed-forward network (FFN) channels. However, balancing model accuracy and inference speed remains challenging. We therefore introduce MAP-MoE, a training-free framework that combines modality-aware expert FFN channel pruning with our Ragged-Hybrid GPU executor. Using a small unlabeled calibration set, MAP-MoE groups experts into visual, mixed, and textual groups and assigns different pruning rates to them under a shared global budget. Within each MoE layer, MAP-MoE ranks FFN channels across experts in the same modality group and retains only the selected rows of the FFN gate and up projections and the corresponding columns of the down projection. Experts retain different numbers of channels after pruning, making standard execution much slower. Ragged-Hybrid addresses this with custom Triton kernels for small token batches and grouped GEMM for larger inputs. Our main evaluation covers two multimodal MoE models at FFN channel pruning ratios of 20%, 30%, and 50% across 10 benchmarks. At 50% pruning, MAP-MoE improves average accuracy by 1.34 points over the SOTA baseline (CAMERA-P) on Qwen3-VL while retaining 98.5% of the original Qwen3-VL average accuracy. For batched generation on one A800, our framework achieves up to 1.65× end-to-end generation speedup over the original model and reduces peak allocated GPU memory by up to 46%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.