The Blind Spot of Post-Training Quantization in Mixture-of-Experts Models
Abstract
In mixture-of-experts (MoE) models, quantization changes which experts each token is routed to. Post-training quantization (PTQ) methods minimize the reconstruction error of each layer, and prior work evaluates the quantized models with benchmark accuracy or the expert-selection change rate. However, we argue that none of these captures how much deviation the routing changes cause. In this paper, we propose a metric for this routing-induced deviation, which compares the quantized model's deviation from the original model with and without fixing its expert selection to the original one. We further derive a law for this deviation, which shows that the layer-wise reconstruction error misses a growing fraction of the deviation as the bit width increases. Lastly, we apply our metric to a released quantized checkpoint and show that it captures a change that accuracy cannot.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.