PaCES: Pattern-Coded Element Sharing for Lightweight Mixture-of-Experts Inference
Abstract
Sparsely activated Mixture-of-Experts (MoE) models improve computational efficiency via token-wise sparse activation, but still require full dense parameter storage, severely limiting their practical deployment. Existing MoE compression approaches fail to achieve a favorable trade-off between model accuracy and storage overhead. In this work, we propose PaCES (Pattern-Coded Element Sharing), a post-training compression method for MoE language models. Based on the key observation that intra-layer experts exhibit consistent Gaussian weight distributions, PaCES groups similar experts and performs distribution-aware element-wise magnitude sharing. It compactly encodes expert participation and sign information using lightweight bit-level codes. Equipped with a fused GPU kernel that avoids dense weight reconstruction during inference, PaCES enables true end-to-end storage and latency optimization. Evaluated on Qwen-MoE and DeepSeek-V2-Lite over eight zero-shot benchmarks, our method reduces checkpoint storage by 32.6%–57.3% and achieves up to 4.05x inference speedup with less accuracy degradation. It consistently outperforms unstructured pruning and expert merging baselines under aggressive compression, providing a practical deployment-friendly compact representation for real-world MoE serving.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.