acceptodds
Under review as a conference paper at ICLR 2027

SpecPilot: Guiding Speculation with Online Adaptive Cost Prediction

Abstract

Speculative decoding speeds up autoregressive inference by having a cheap pro- poser draft several tokens that the target model then verifies in a single forward pass. Whether a step pays off depends on how far acceptance holds, what verify- ing a given tree costs on the machine at hand, and what drafting itself costs, and all three change during a run. Fixed policies therefore win on some workloads and lose on others: the best draft length for llama.cpp’s speculative decoding is the worst on prose, where it runs 0.69× as fast as not speculating at all. We present SpecPilot, a planner that decides at every step how many drafter levels to buy and which tree to verify, judging each choice against the throughput the run has actually achieved. SpecPilot learns the verify cost surface online, including kernel-switch discontinuities whose direction differs between GPUs, and can pad a tree past one when the larger tree is cheaper. It prices each node by a path probability calibrated against observed acceptance, and requests another drafter level only while its expected yield covers the measured cost of drafting. With a single untuned configuration on an RTX 3090, SpecPilot reaches a 1.78× geometric-mean speedup over plain decoding across five workloads, 14.1% faster than llama.cpp's builtin speculative decoding and 3.7% faster than CAST, each tuned per cell in hindsight. CAST’s best thresh- old changes with the workload and the GPU: across 23 cells on three GPUs, SpecPilot matches CAST’s best fixed threshold (+0.5%) and is within 0.3% of CAST retuned for every cell, while CAST’s worst threshold falls up to 22.7% below it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.