acceptodds
Under review as a conference paper at ICLR 2027

HAMoE: Heterogeneity-Aware Mixed-Precision Quantization for Mixture-of-Experts Models

Abstract

Mixture-of-Experts (MoE) models face deployment challenges in resource-constrained devices due to their massive parameter scales. Quantization is recognized as one of the most effective approaches to address such challenges. In this work, we conduct a systematic quantization study of MoE models and identify two critical heterogeneity insights that have not been considered in prior work: 1) inter-expert heterogeneity, i.e., distinct activation frequencies and processed token types across experts, and 2) intra-expert heterogeneity, i.e., differential quantization sensitivity of weights within a single expert. To this end, we introduce HAMoE, a mixed-precision quantization framework that considers both heterogeneities. HAMoE assigns fine-grained mixed precision to inter-experts based on activation frequency and token type. In addition, it proposes an adaptive column replacement mechanism, which replaces the most quantization-sensitive intra-expert weight columns with INT8-quantized versions, thereby effectively reducing quantization loss. Evaluation results show that HAMoE achieves competitive results against strong baselines under equivalent bit budgets, reducing WikiText-2 perplexity by 4.0% relative to GPTQ at 3.35-bit, recovering about 45% of the quantization loss. HAMoE also consistently outperforms MxMoE in both average accuracy and perplexity across all four models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.