HeteroQuant: Resolving Expert Heterogeneity in MoE Quantization via Targeted Calibration
Abstract
Mixture-of-Experts (MoE) architectures substantially expand model parameter capacity while maintaining a limited computational cost per token through sparse activation and dynamic routing mechanisms. However, during practical deployment, the complete set of expert parameters usually needs to remain in memory, and the resulting large storage overhead restricts the application of MoE models in resource-constrained environments. Post-training quantization (PTQ) is an important technique for reducing the storage and inference costs of large language models. Nevertheless, due to the significant expert heterogeneity, domain preferences, and imbalanced activation patterns in MoE models, existing quantization methods often fail to sufficiently calibrate critical experts, leading to severe performance degradation. To address these challenges, we propose HeteroQuant, an effective, training-free post-training quantization framework designed to resolve expert heterogeneity through targeted calibration. Specifically, HeteroQuant achieves this by first leveraging the DSPy framework to construct a highly diverse multi-domain candidate pool, thereby comprehensively activating all expert positions to collect their global routing statistics. Based on these statistics, we implement a core two-step selection mechanism to systematically optimize the calibration data. In the first stage, we perform AlphaQ-guided structural-targeted scoring to prioritize candidate samples that precisely target and activate structurally critical experts. In the second stage, we execute activation-redundancy-aware filtering to eliminate samples with overlapping routing footprints, thereby preventing calibration waste while maximizing activation diversity. Experimental results demonstrate that under the W4A4 quantization setting of the Qwen1.5-MoE-A2.7B model, HeteroQuant achieves an average improvement of 5.09 points over the current SOTA method across eight downstream tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.