acceptodds
Under review as a conference paper at ICLR 2027

SpecRouter: Reusable Payoff Prediction for Cost-Aware Speculative Decoding

Abstract

Speculative decoding can obtain draft tokens in very different ways—copying from the context, n-gram lookup, a small draft model, or a learned tree drafter—and which drafter is fastest depends both on the text being generated and on the hardware and batch size that execute it. We introduce , which separates these two factors. A 0.58M-parameter predictor reads the target model's hidden state and estimates how many tokens each drafter would commit; a separately calibrated cost model prices each drafter on the current hardware; only the drafter with the lowest predicted cost per committed token runs. The predictor needs no exploration to train: with deterministic drafters and greedy verification, the output is the same whichever drafter runs, so every drafter's committed length at every prefix can be computed exactly offline. We evaluate on sixteen workloads with 100 held-out requests each. Averaged over workloads, speeds up decoding by on Llama-3.1-8B and on Qwen3-8B, which is 30% and 27% higher than the speedup of EAGLE-3, the strongest fixed drafter. It also exceeds the best fixed drafter chosen separately for each workload in hindsight. Comparing costs explicitly is the largest single source of this gain. also outperforms online methods that learn drafter choice during generation, reaching versus on a pool of domain-specialized drafters. A single predictor, never retrained, is the fastest policy on A40 and H200 GPUs at batch sizes 1, 4, and 16; only its cost model is recalibrated.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.