acceptodds
Under review as a conference paper at ICLR 2027

MoE-Spec: Expert Budgeting for MoE Speculative Decoding

Abstract

Speculative decoding accelerates Large Language Model (LLM) inference by verifying multiple candidate tokens in parallel. For Mixture-of-Experts (MoE) models, each candidate token routes to different experts, and the target model must load the union of all activated experts to verify them. As the number of candidates grows, unique experts per layer approach the full model, and verification becomes as expensive as running a dense model. We propose MoE-Spec, a training-free method that enforces a fixed expert budget at each layer during verification. MoE-Spec aggregates routing probabilities across candidate tokens and loads only the top- experts. Because the remaining experts are excluded, the budget controls a tradeoff between throughput and quality. At conservative budgets, MoE-Spec matches EAGLE quality and often improves throughput; tighter budgets trade quality for additional throughput. Experiments on OLMoE-1B-7B, Qwen3-30B-A3B, and Mixtral-8x7B across reasoning, code generation, and summarization show that at conservative budgets MoE-Spec matches EAGLE quality while improving throughput, and at aggressive budgets retaining 90% of autoregressive quality, throughput improves 15–45% depending on model architecture. Code will be released upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.