Window Pool: Verify-Aligned Routing for MoE Speculative Decoding
Abstract
Speculative decoding (SD) verifies draft tokens in a single target forward pass. For a Mixture-of-Experts (MoE) target this parallelism undercuts the intended saving: the verification block activates the union of the per-token top- expert sets, so expert-weight traffic grows even though the token count is amortized. Existing remedies cap experts at serve time or reshape the draft tree, but leave the pretrained router untouched. We instead attack the union inflation at its source: the pretrained routing prior. We first show that document-level expert-pool pretraining already induces a prior favorable to speculation—lower verification inflation, and graceful rather than catastrophic degradation under forced window sharing—and then build on it. We propose Window Pool, a continued-pretraining recipe that aligns the expert-sharing group with the SD verification window, with randomized window length and pool size, layer selection, and interleaved unconstrained steps. Our main result is on the deployment grid: across a sweep of prefill-only window-sharing constraints, Window Pool beats the document-pool base in every one of the twelve constrained cells (typical gap – mean Avg), while preserving unconstrained quality. When the same routers are placed under a hard per-verify-round expert budget (uniform MoE-Spec substitution on all MoE layers), the trained prior—not the serve-time cap—is what makes the budget survivable: at the standard router collapses (MMLU , LAMBADA ) while Window Pool retains MMLU and LAMBADA . On HumanEval-164 end-to-end SD, Window Pool does not regress under budgeting (, at traffic ) where the standard router falls to / ; against the document-pool base the SD gap is within task-count noise, so the contribution is constrained-sharing robustness rather than SD speedup. Window Pool is a training-time routing prior rather than a serve-time heuristic: by internalizing the verification constraint during continued pretraining, it renders aggressive expert-budgeting viable for MoE speculative decoding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.