Fragility-Aware Adaptive Routing for Quantized Mixture-of-Experts Serving
Abstract
Quantizing the experts of a Mixture-of-Experts (MoE) model is standard practice for relieving memory pressure during serving, producing a pool of pre-quantized model copies: instances, whose quality cost varies sharply across requests. Existing systems have two limitations: statically committing to a single instance for the workload and relying on ad hoc heuristic request score predictors. Closing this gap, we propose an adaptive routing framework with theoretical guarantees that assigns each incoming request to an instance so as to maximize throughput within a quality-degradation budget. Underlying the framework is a structural result: a small-noise expansion factors quality loss over experts into request-independent expert fragility and request-dependent expert affinity. This factorization yields FWP (Fragility-Weighted Perplexity), a lightweight quality estimator read from a single prefill pass on the cheapest instance and transferred to every other instance by affine calibration. Building on FWP, we cast routing as a window-level linear program and prove that the induced per-request greedy policy is KKT-consistent with the LP optimum. Empirically, on complete Qwen3-30B-A3B instances, FWP- guided allocation at a fixed quality budget comes within % of an allocation that knows each request’s true quality loss, well ahead of activation-frequency and prompt-length signals. Turning this signal into live serving, the LP-derived router over 4-bit, 4/8-bit mixed, and 8-bit instances meets every per-class quality budget at up to the measured throughput of the only static deployment that satisfies all classes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.