acceptodds
Under review as a conference paper at ICLR 2027

FairSpec: Robust Expert Selection for Efficient MoE Speculative Decoding

Abstract

Speculative decoding accelerates large language-model (LLM) inference by using a lightweight drafter model to propose several tokens and verifying them simultaneously in one target-model pass. For mixture-of-experts (MoE) models, however, the speculative tokens may route to different experts, forcing each MoE layer to read more distinct expert parameters from GPU high-bandwidth memory and potentially making verification slower than autoregressive decoding. Existing methods cap this cost by retaining the experts with the greatest aggregate routing demand from speculative tokens, but aggregate selection can neglect individual tokens, shortening the accepted draft prefix and/or degrading model inference quality. We therefore propose FairSpec, a verifier-side expert routing method that allocates an expert budget that protects each speculated token's highest-ranked experts by exploiting expert overlap across tokens. FairSpec requires no model retraining or changes to the downstream MoE operator and is implemented as a fused, CUDA-graph-compatible routing kernel. We evaluate FairSpec against state-of-the-art speculative decoding methods that perform expert-subset selection. Across 4 benchmarks and 3 sparse MoE architectures, at comparable throughput, FairSpec achieves 10–50pp higher success rates under low and moderate expert budgets, with larger gains at lower budgets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.