SLaMES: Sensitivity-Guided Layerwise and Modality-Aware Expert Skipping for Efficient Inference in MoE MLLMs
Abstract
Mixture-of-Experts (MoE) multimodal large language models (MLLMs) scale capacity efficiently, but fixed Top- routing activates the same number of experts per token, causing substantial inference overhead for visual-token-heavy inputs. Existing training-free expert-skipping methods use coarse, layer-shared, or skip-rate-invariant criteria, overlooking that skipping tolerance varies across layers, modalities, and ratios. We propose SLaMES, a training-free adaptive framework that exploits this heterogeneity through three components. First, *Layer–Modality Sensitivity Calibration* quantifies sensitivity across layers, modalities, and skipping ratios. Second, *Sensitivity-Guided Threshold Search* uses calibrated sensitivities and offline lookup to derive layer- and modality-specific thresholds with only a few full-model forward passes. Third, *Adaptive Routing-Based Skipping* allocates expert computation based on cumulative routing probability mass, with conditional Top-1 protection for the dominant expert. Across three MoE MLLMs and nine image and video benchmarks, SLaMES retains at least 99.25% of the original performance at a 50% expert-skipping ratio. At the most aggressive settings of 83.4%–87.5%, it retains 96.39%–98.14% of the original performance, consistently outperforming existing expert-skipping methods. Integrated into vLLM, these savings translate into prefill speedups of up to 2.2 on real hardware, demonstrating the efficiency of SLaMES.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.