MoE-VQ: A Model-Agnostic Gaussian Vector Codebook for 2-Bit Expert Weight Quantization
Abstract
Mixture-of-Experts (MoE) models concentrate most parameters in sparsely activated experts, making weight memory rather than computation the primary deployment bottleneck. However, scalar post-training quantization incurs severe and irreversible distortion at 2 bits. We show that, after row-group RMS normalization, expert MLP weights from six MoE architectures closely follow a shared standard Gaussian distribution. This observation enables a single, model-agnostic vector-quantization codebook: we fit an 8-dimensional, 2^16-entry codebook once to N(0, I) and reuse it across models, layers, and experts without model-specific codebook training. Building on this codebook, MoE-VQ combines block-wise LDLQ residual feedback with rank-14 low-rank error compensation. At 2.34 effective bits per weight, MoE-VQ improves average accuracy across seven tasks by 3.3–13.3 points over tuned 2-bit GPTQ on four MoE models, while removing 51% of the excess perplexity relative to FP16. It also better preserves expert routing and outperforms an equal-rate E8 lattice by 1.0–3.2% in perplexity. For Mixtral-8x7B, MoE-VQ reduces weight memory from 93.4 GB to 15.6 GB, allowing the validated weight configuration to fit on one 80 GB GPU instead of two. These results show that distribution-adaptive vector codebooks can be reusable rather than model-specific for low-bit MoE quantization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.