SAGE-VQ: Structure-Adaptive Gating-Aware Expert Vector Quantization with Output Correction for MoE LLMs
Abstract
Mixture-of-Experts (MoE) large language models (LLMs) expand capacity through sparse activation, yet expert storage remains a bottleneck for single-GPU and on-device deployment. Post-training vector quantization (VQ) enables ultra-low-bit compression, but dense-LLM methods overlook heterogeneous expert sensitivity, architecture-dependent output composition, and routing-weighted errors. We introduce SAGE-VQ, a structure-adaptive, gating-aware expert VQ framework that aligns precision allocation and output correction with MoE structure. SAGE-VQ selects architecture-appropriate sensitivity signals to assign expert bit-widths under a fixed average bit budget. It incorporates routing weights into the correction objective to mitigate norm and scale mismatches in quantized expert outputs. Correction factors are fused before inference without additional runtime latency. On Qwen2-MoE, SAGE-VQ achieves 2.25-bit expert compression, reducing WikiText-2 perplexity by 1.21 relative to the MoE-aware mixed-precision baseline MxMoE while limiting average accuracy degradation across seven zero-shot benchmarks to 1.42 percentage points relative to FP16.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.