AdaSpec: Speculative Decoding with In-Engine Continual Learning
Abstract
Speculative decoding accelerates large language model inference with a small drafter. However, the drafter is frozen after offline training, while it serves narrow and shifting traffic on which its predictions degrade. Recent work trains the drafter online outside the inference engine, which costs extra GPUs and data movement and serves with stale weights. We introduce AdaSpec, an inference engine in which the drafter learns continually from the traffic it serves, on the same GPUs and from the same forward passes. AdaSpec makes learning a by-product of serving: it reuses the activations and verification results already on the GPU and overlaps the backward pass with decoding. The drafter learns to maximize the expected number of accepted tokens, and a cost-aware gate trains it only while the gain pays for the cost. On streams that mimic real deployments, from agents to low-resource languages and real users, AdaSpec improves the acceptance length by 10.6% and 14.7% on average at two model scales, and the end-to-end throughput by 4.3% and 4.7%, including all learning costs. These results indicate that an inference server can become faster the longer it serves.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.