acceptodds
Under review as a conference paper at ICLR 2027

Lost in Pruning: Silent Multilingual Collapse and Training-Free Recovery for One-Shot MoE Pruning

Abstract

One-shot expert pruning shrinks the resident memory of sparse mixture-of-experts (MoE) models without retraining and is validated on benchmarks that resemble its calibration data. We show that this validation misses a silent multilingual collapse: a 25%-pruned Qwen3-30B-A3B keeps code near baseline and English perplexity within ×1.6, yet answers Korean prompts in Chinese. The damage is a loss of language access rather than of knowledge alone: on 2,800 parallel MMLU-ProX items the pruned model loses 56 points in Korean but 15 in English, and still answers 74% of its Korean losses correctly in English. Fluency, knowledge, answer extraction, and translation direction fail differently; on this model generative MMLU-ProX also falls to near zero in Japanese and French, generation into an uncovered language degrades far more than translation out of it, and the collapse recurs across six MoE architectures and four scoring criteria, implicating calibration coverage rather than any single scoring rule. Controlled expert exchanges at fixed compression, each designed before its checkpoints were built, probe this account directly: restoring 28% of the experts on which the default mask and a multi-view coverage mask disagree closes about 80% of the Korean NLL gap between the two masks and restores Korean-script responses, but only 11% of the KMMLU gap. This language-before-knowledge ordering recurs on a 256-expert model, where it was predicted before the run, and is absent in a model whose default pruning does not collapse Korean generation. Building on this analysis, Cross-Observation Expert Retention (COER) recomposes cached per-domain calibration views on CPU without raw text or retraining; at matched compression and observation budgets (three seeds on Qwen3) it restores target-language generation under all four criteria, generalizes to five languages without labeled calibration examples in those languages, and recovers 34–77% of lost HAERAE accuracy under gate weighting, with rank-union retaining code at or above the Qwen3 default.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.