Rep-Q: Replaceability-Aware Mixed-Precision Quantization for MoE LLMs
Abstract
Mixture-of-Experts Large Language Models (MoE-LLMs) can expand model capacity while controlling computational costs, yet their massive expert parameters still introduce substantial memory overhead. Mixed-precision quantization mitigates this issue by assigning differentiated bit widths to distinct experts. Existing approaches mostly assess expert importance based on activation frequency or loss sensitivity but largely overlook expert replaceability, i.e., whether the functionality of one expert can be fulfilled by others. Highly replaceable experts encode knowledge already covered by other experts, but experts with low replaceability require higher precision. Accordingly, this paper proposes Rep-Q, a mixed-precision quantization framework that jointly models expert replaceability and importance for quantization bit-width adjustment. (1) To translate the directly unmeasurable replaceability among experts into a computable metric, we design a similarity score to quantify similarities between expert output representations, establishing a quantitative foundation for modeling information dependencies across experts. (2) Leveraging these similarity scores, we construct a unified expert uniqueness metric by aggregating representation similarities between each expert and all remaining experts. This metric identifies single-expert replaceability from pairwise expert correlations and quantifies the irreplaceability of each expert’s information, enabling the transition from relationship-level measurements to individual-level evaluations. (3) Subject to quantization budget constraints, we jointly model expert quantization sensitivity and uniqueness scores. This allocation ensures experts with high importance and low replaceability are assigned higher precision. Evaluations across multiple datasets on Mixtral 8×7B (pure routing experts) and DeepSeek-V2-Lite (with shared experts) demonstrate that under an identical quantization budget, Rep-Q yields consistent improvements in perplexity and downstream task accuracy over existing expert-level mixed-precision quantization baselines, and exhibits strong cross-architecture generalization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.