D-cut: Adaptive Verification Depth Pruning for Batched Speculative Decoding
Abstract
Speculative decoding has become a leading approach for accelerating large language model (LLM) inference without compromising output quality. Recent methods further improve single-request speedup by decoupling draft length from drafting latency via parallel generation, which produces longer drafts to achieve higher mean accepted tokens (MAT). However, as request concurrency grows, these long drafts waste more compute on rejected tokens, inflating the *verification cost* and making speculative decoding suboptimal in high-concurrency scenarios. To address this challenge, we propose D-cut, an adaptive pruning strategy that selects draft tokens across the entire batch, concentrating verification budget on tokens most likely to be accepted. D-cut is motivated by two observations: (1) acceptance length varies significantly across concurrent requests, so D-cut performs *cross-request pruning* that adaptively allocates verification budget based on draft confidence; (2) verification cost largely depends on the runtime environment (GPU, parallelism, etc.), so we incorporate a *runtime cost model* into the selection, allowing D-cut to adapt its pruning depth to the actual deployment. Experiments across dense and MoE models show that D-cut improves the average speedup from 1.26× to 1.65× under high concurrency, restores acceleration for dense-model configurations where long-draft baselines fall below autoregressive decoding, and achieves up to 3.0× speedup over autoregressive decoding on MoE models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.