Static Preservation Is Not Enough: Adaptation-Aware Structural Search for MoE Pruning
Abstract
Mixture-of-Experts (MoE) language models activate only a few experts per token but retain the full expert bank in memory during inference. Expert pruning reduces this deployment overhead by removing experts, but existing methods commonly evaluate retention decisions using layer-local proxies or a fixed model state. Such estimates can be unreliable because retention decisions interact across layers, while structures that preserve the current model well may not remain effective after subsequent adaptation. We therefore introduce MASS, a Model-Wide Adaptation-Aware Structural Search method, in which retention decisions remain revisable as the model evolves. Specifically, we evaluate candidate structures jointly across all MoE layers using full-model next-token prediction feedback. Under the target expert budget, a routing-aware objective guides expert allocation according to token-level native routing demand without favoring capacity expansion. We couple structural search with model adaptation throughout optimization, where current retention preferences define the structure used for model updates and feedback from the adapted model guides their subsequent refinement. Experiments on two distinct MoE architectures under two pruning severities show that MASS achieves the best overall recovered performance. Further analyses show that MASS exhibits favorable representation recovery while preserving native routing demand and limiting post-pruning representation distortion.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.