acceptodds
Under review as a conference paper at ICLR 2027

SAGE-VQ: Structure-Adaptive Gating-Aware Expert Vector Quantization with Output Correction for MoE LLMs

Abstract

Mixture-of-Experts (MoE) large language models (LLMs) expand capacity through sparse activation, yet expert storage remains a bottleneck for single-GPU and on-device deployment. Post-training vector quantization (VQ) enables ultra-low-bit compression, but dense-LLM methods overlook heterogeneous expert sensitivity, architecture-dependent output composition, and routing-weighted errors. We introduce SAGE-VQ, a structure-adaptive, gating-aware expert VQ framework that aligns precision allocation and output correction with MoE structure. SAGE-VQ selects architecture-appropriate sensitivity signals to assign expert bit-widths under a fixed average bit budget. It incorporates routing weights into the correction objective to mitigate norm and scale mismatches in quantized expert outputs. Correction factors are fused before inference without additional runtime latency. On Qwen2-MoE, SAGE-VQ achieves 2.25-bit expert compression, reducing WikiText-2 perplexity by 1.21 relative to the MoE-aware mixed-precision baseline MxMoE while limiting average accuracy degradation across seven zero-shot benchmarks to 1.42 percentage points relative to FP16.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.