VerifyLocal: Learning Expert-Residency-Conditioned Proposals for Exact MoE Speculative Decoding
Abstract
Exact speculative decoding fixes the returned target distribution but permits its auxiliary proposal to use observable deployment state. In expert-offloaded mixture-of-experts (MoE) inference, plausible speculative histories can incur different expert-weight transfers. VerifyLocal conditions proposal probabilities on pre-round expert residency while keeping the target, its router, and exact correction unchanged. The Stage-A screen reduces logical miss bytes by 30% (95% CI [24%, 36%]) while retaining 98.5% expected commitment. In the controlled physical evaluation, expert-transfer bytes fall from 84.0 to 63.0 MB per committed token (25%; 95% CI [20%, 30%] reduction), and committed-token goodput rises from 19.5 to 23.0 tokens/s (18%; 95% CI [13%, 23%] increase). Inter-token latency falls by 18% and time to first token by 10%; observed task quality differs by -0.1 percentage points (95% CI [-0.3, +0.1]). Goodput exceeds a tuned nonlearned residency-aware cost tilt by 8% (95% CI [+4%, +12%]). Physical-transfer reductions replicate on OLMoE-1B-7B (27%, 95% CI [21%, 33%]), Mixtral-8x7B (24%, [17%, 30%]), and DeepSeekMoE-16B (22%, [15%, 29%]). The benefit diminishes toward zero with fully resident experts, delimiting the movement-limited regime.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.