CORTEX: COmpressing RouTed EXperts for MoE-based LLMs under the Block Loss
Abstract
Mixture-of-Experts (MoEs) scale the sizes of LLMs, thus studying their compression becomes crucial. One distinctive feature of MoEs is that the experts of a layer are redundant in their weights and correlated in their routing: a token activates several of them at once, and the subsequent layer receives only their gated sum. Existing low-rank compression methods exploit the redundancy through parameters shared across experts, but the routing correlation rarely enters their objectives: each expert is fitted on its own inputs. In this work, we propose to fit the low-rank compression under the error of the routed MoE block, in which the errors of co-routed experts are coupled. The exact second-order metric of this loss, however, couples all experts and admits no closed-form truncation. Besides, its per-expert blocks are estimated from few tokens and overfit the calibration set. Our method therefore uses two surrogates of this metric, one keeping the co-routing interference and the other the per-expert input covariances. Under the surrogate metrics with a Stein-type shrinkage, each component of the low-rank decomposition is fitted in a closed form, and the output factors of all experts are then re-solved jointly under the exact metric. Finally, a role-split rank allocation from a first-order analysis of the block output completes the method. We demonstrate the method at –% compression on Mixtral-87B, Qwen2-57B-A14B, and DeepSeekMoE-16B-Base, where it achieves the strongest performance in language modeling and zero-shot reasoning in most settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.