BLAZE: Bias-Driven Load-Aware Zero-Overhead Expert Routing for MoE LLM Inference
Abstract
Expert-parallel Mixture-of-Experts (MoE) inference is usually bottlenecked by routing-induced load imbalance: overloaded experts and EP ranks create stragglers in All-to-All communication and grouped-GEMM execution. Existing placement-level approaches mitigate persistent hotspots by replicating expert weights, but this consumes GPU memory and does not directly correct transient router imbalance at inference time. We present BLAZE, a token-drop-free and replica-free router-side controller for expert-parallel MoE serving. BLAZE redirects only confidence-ambiguous expert choices away from overloaded experts or EP ranks, while locking high-confidence semantic decisions and computing mixture weights from the original router scores. To make redirection deployable in end-to-end serving, BLAZE further uses an overhead-aware activation gate that bypasses the controller on small effective MoE batches where redirection overhead is unlikely to pay off. Across GPT-OSS-20B and Qwen3-30B-A3B, BLAZE stays within 1.6 percentage points of vanilla accuracy, while random or worst-ranked replacement collapses accuracy by 44–49 percentage points; on GPT-OSS-120B, three complete MLPerf accuracy runs of BLAZE all pass the official threshold with no loss relative to vanilla. On DeepSeek-R1, the overhead-aware gate avoids a 16% TPOT regression from always-on redirection. Under MLPerf-style serving with TP = EP = 8 at SLO-feasible load, BLAZE turns a vanilla TTFT violation on GPT-OSS-20B into an SLO pass (26.7% lower TTFT P99) and, in a matched QPS sweep, still meets the latency SLO at 1.75 QPS, whereas capacity-aware routing fails from 1.25 QPS; on Qwen3-30B-A3B, all methods meet the SLO and BLAZE delivers the highest completed-token throughput. BLAZE also improves offline throughput by 8.9% on GPT-OSS-120B.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.