acceptodds
Under review as a conference paper at ICLR 2027

GLoRE: Global-Local Router-Mass Expert Skipping for Batched MoE Serving

Abstract

Mixture-of-Experts (MoE) scales model capacity without proportionally increasing per-token compute and now underpins many large language models. This sparsity is less effective during serving: decode is dominated by weight movement, so each step's cost depends on the batch's union of selected experts rather than any one token's top- experts. We introduce GLoRE, a runtime expert-masking method that reduces this union through global expert masking for efficiency and local protection for quality. At each decoding step, GLoRE masks experts with low aggregate router gate mass under a batch-level budget while protecting a token's top expert when its routing weight exceeds a threshold. This adapts expert reduction to each batch while preserving important token-level assignments, and remains robust as batch sizes vary. Implemented as a plug-and-play CUDA kernel for vLLM, GLoRE achieves up to decode throughput at batch size 64 with only 2% accuracy loss on Qwen3-30B-A3B, and remains on the accuracy–throughput Pareto frontier across four MoE models. On long chain-of-thought test-time scaling, GLoRE maintains vanilla-level maj@8 accuracy on AIME and GPQA-Diamond while using up to fewer GPU-seconds.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.