acceptodds
Under review as a conference paper at ICLR 2027

AdaSpec: Speculative Decoding with In-Engine Continual Learning

Abstract

Speculative decoding accelerates large language model inference with a small drafter. However, the drafter is frozen after offline training, while it serves narrow and shifting traffic on which its predictions degrade. Recent work trains the drafter online outside the inference engine, which costs extra GPUs and data movement and serves with stale weights. We introduce AdaSpec, an inference engine in which the drafter learns continually from the traffic it serves, on the same GPUs and from the same forward passes. AdaSpec makes learning a by-product of serving: it reuses the activations and verification results already on the GPU and overlaps the backward pass with decoding. The drafter learns to maximize the expected number of accepted tokens, and a cost-aware gate trains it only while the gain pays for the cost. On streams that mimic real deployments, from agents to low-resource languages and real users, AdaSpec improves the acceptance length by 10.6% and 14.7% on average at two model scales, and the end-to-end throughput by 4.3% and 4.7%, including all learning costs. These results indicate that an inference server can become faster the longer it serves.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.