AdaptSpec: Input- and Step-Adaptive Tree-Based Speculative Decoding to Accelerate LLM Inference
Abstract
Speculative decoding accelerates Large Language Model (LLM) inference by using a draft model to propose a speculative length (SL) of tokens at each decoding step, which a target model then verifies in a single forward pass. Tree-based speculative decoding further improves acceptance rates and speedup by drafting a candidate length (CL) of tokens in parallel at each position. However, these methods rely on fixed SL and CL. While recent work adapts SL primarily using step-level signals, prior methods do not jointly optimize SL and CL while conditioning on the request input, despite its significant impact on optimal parameter selection. Our experimental analysis identifies several influential factors, including input categories, that determine the speedup-maximizing configuration. Accordingly, we propose AdaptSpec, which initially adapts SL using machine learning and CL using binary search based on these factors, and subsequently switches to Reinforcement Learning (RL) to jointly optimize both parameters once the RL policy is trained. Experiments across four draft/target model pairs and five datasets show that AdaptSpec achieves up to 98% higher speedup than state-of-the-art methods, translating into fewer GPU-hours per request in LLM serving.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.