GLoRE: Global-Local Router-Mass Expert Skipping for Batched MoE Serving
Abstract
Mixture-of-Experts (MoE) scales model capacity without proportionally increasing per-token compute and now underpins many large language models. This sparsity is less effective during serving: decode is dominated by weight movement, so each step's cost depends on the batch's union of selected experts rather than any one token's top- experts. We introduce GLoRE, a runtime expert-masking method that reduces this union through global expert masking for efficiency and local protection for quality. At each decoding step, GLoRE masks experts with low aggregate router gate mass under a batch-level budget while protecting a token's top expert when its routing weight exceeds a threshold. This adapts expert reduction to each batch while preserving important token-level assignments, and remains robust as batch sizes vary. Implemented as a plug-and-play CUDA kernel for vLLM, GLoRE achieves up to decode throughput at batch size 64 with only 2% accuracy loss on Qwen3-30B-A3B, and remains on the accuracy–throughput Pareto frontier across four MoE models. On long chain-of-thought test-time scaling, GLoRE maintains vanilla-level maj@8 accuracy on AIME and GPQA-Diamond while using up to fewer GPU-seconds.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.