acceptodds
Under review as a conference paper at ICLR 2027

PaCES: Pattern-Coded Element Sharing for Lightweight Mixture-of-Experts Inference

Abstract

Sparsely activated Mixture-of-Experts (MoE) models improve computational efficiency via token-wise sparse activation, but still require full dense parameter storage, severely limiting their practical deployment. Existing MoE compression approaches fail to achieve a favorable trade-off between model accuracy and storage overhead. In this work, we propose PaCES (Pattern-Coded Element Sharing), a post-training compression method for MoE language models. Based on the key observation that intra-layer experts exhibit consistent Gaussian weight distributions, PaCES groups similar experts and performs distribution-aware element-wise magnitude sharing. It compactly encodes expert participation and sign information using lightweight bit-level codes. Equipped with a fused GPU kernel that avoids dense weight reconstruction during inference, PaCES enables true end-to-end storage and latency optimization. Evaluated on Qwen-MoE and DeepSeek-V2-Lite over eight zero-shot benchmarks, our method reduces checkpoint storage by 32.6%–57.3% and achieves up to 4.05x inference speedup with less accuracy degradation. It consistently outperforms unstructured pruning and expert merging baselines under aggressive compression, providing a practical deployment-friendly compact representation for real-world MoE serving.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.