acceptodds
Under review as a conference paper at ICLR 2027

MoE-VQ: A Model-Agnostic Gaussian Vector Codebook for 2-Bit Expert Weight Quantization

Abstract

Mixture-of-Experts (MoE) models concentrate most parameters in sparsely activated experts, making weight memory rather than computation the primary deployment bottleneck. However, scalar post-training quantization incurs severe and irreversible distortion at 2 bits. We show that, after row-group RMS normalization, expert MLP weights from six MoE architectures closely follow a shared standard Gaussian distribution. This observation enables a single, model-agnostic vector-quantization codebook: we fit an 8-dimensional, 2^16-entry codebook once to N(0, I) and reuse it across models, layers, and experts without model-specific codebook training. Building on this codebook, MoE-VQ combines block-wise LDLQ residual feedback with rank-14 low-rank error compensation. At 2.34 effective bits per weight, MoE-VQ improves average accuracy across seven tasks by 3.3–13.3 points over tuned 2-bit GPTQ on four MoE models, while removing 51% of the excess perplexity relative to FP16. It also better preserves expert routing and outperforms an equal-rate E8 lattice by 1.0–3.2% in perplexity. For Mixtral-8x7B, MoE-VQ reduces weight memory from 93.4 GB to 15.6 GB, allowing the validated weight configuration to fit on one 80 GB GPU instead of two. These results show that distribution-adaptive vector codebooks can be reusable rather than model-specific for low-bit MoE quantization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.