Owen-Value-Guided Mixed-Precision Quantization for Mixture-of-Experts Models
Abstract
Mixture-of-experts (MoE) models activate only a few experts per token, yet low-latency inference can require the entire expert pool in memory. At ultra-low average precision, mixed-precision quantization must decide which experts retain scarce higher-bit slots. Individual sensitivity scores can miss the joint effects of co-routed experts. We propose Owen-MoE, which values experts in an availability game while accounting for an output-similarity partition. Two-level Monte Carlo sampling estimates Owen values in O(TK) evaluations per layer for T trials over K experts. A scale-invariant Influence Ratio converts these values into a static 2,4-bit map that can protect locally influential experts within lower-value coalitions. We compare rankings under identical per-layer 4-bit counts and AutoRound settings. Across eight zero-shot benchmarks on Qwen3-30B-A3B, three-seed means at 2.75–2.125 average expert bits, Owen-MoE exceeds Router-norm by 3.57–7.02 points and PMQ by 4.86–10.56 points on average. It also exceeds measured fractional Uniform by 2.78–13.31 points at matched expert-weight payload. Equal-budget controls show an additional gain over flat Shapley of 1.78 points at 2.50 bits and 2.58 points at 2.25 bits. In 96 selected high-coactivation expert pairs where the rankings disagree, Owen's choice reduces round-to-nearest quantization MSE in 78.1% of cases, with a median reduction of 17.6%. Mixtral and Qwen1.5-MoE extend the evaluation to different expert counts and routing configurations. We will release the complete code upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.