acceptodds
Under review as a conference paper at ICLR 2027

MxQBound: Efficient Mixed-Precision Quantization for MoE-LLMs Guided by Expert Importance and Error Bound Guarantees

Abstract

While mixed-precision quantization reduces the memory overhead of Mixture-of-Experts (MoE)-Large Language Models (LLMs), existing methods often mis-estimate expert importance, causing suboptimal precision allocation and severe performance degradation under aggressive quantization. We propose MxQBound, an efficient and robust Mixed-precision Quantization framework for MoE-LLMs, guided by holistic assessment of expert importance and equipped with quantization error Bound guarantees. With only a single forward pass, MxQBound directly assesses each expert’s importance from its output norms, supported by a provable upper bound on quantization error. Guided by the importance assessment, MxQBound jointly optimizes the per-expert bit-widths of weights and activations to minimize quantization errors under an overall bit-allocation budget. It further adapts each expert’s activation bit-width according to its input strength during inference with negligible overhead. Extensive experiments on diverse MoE-LLMs and benchmarks show that, under comparable compression ratios, MxQBound outperforms state-of-the-art (SOTA) methods by 4.08-22.85% in model performance and 2.2-141× in computational efficiency. Under aggressive W4A4, MxQBound incurs an average accuracy drop of only 15.07% relative to the full-precision baseline, whereas SOTA methods collapse to near-zero accuracy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.