Beyond Static Importance: Noise-Aware Expert Pruning for Diffusion Language Models
Abstract
Expert pruning has been widely studied for autoregressive Mixture-of-Experts (MoE) language models as an effective way to reduce memory overhead by removing less important experts. However, pruning methods developed for autoregressive models do not transfer well to diffusion language models (dLLMs), which perform multi-step denoising and may rely on different experts across noise levels, unlike autoregressive models that generate text along a single left-to-right trajectory. Our analysis shows that routing and activation statistics can mis-rank experts under the denoising objective, while expert importance varies substantially across noise levels, with different rankings at high and low noise levels. Consequently, evaluating experts under a single input distribution or fixed noise level may retain redundant experts while pruning those needed at other denoising stages. Motivated by these observations, we propose NAPE, a one-shot pruning criterion that scores each expert using the absolute gradient of the masked-prediction loss with respect to its gate and aggregates the score over mask ratios spanning high- to low-noise stages. By aggregating expert importance across noise levels, NAPE avoids biased estimates from calibration at fixed noise. Experiments on two MoE dLLM backbones across ten downstream benchmarks show that NAPE consistently outperforms existing pruning methods under different pruning ratios, improving the strongest baseline by up to 12.3% relative at 50% sparsity. We release our code at https://anonymous.4open.science/r/NAPE-35C2 to facilitate future research on efficient MoE dLLMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.