acceptodds
Under review as a conference paper at ICLR 2027

Training-Free Dynamic Verification Budget for Parallel Speculative Decoding

Abstract

Parallel speculative decoding reduces drafting latency through one-shot block generation, but its end-to-end efficiency increasingly depends on how target-model verification is budgeted. A fixed verification budget wastes computation on difficult prompts with limited acceptance potential while restricting acceleration on easy ones. We characterize this mismatch through the trade-off between affine verification costs and strict diminishing marginal acceptance gains, yielding a unique interior optimum that varies dynamically with prompt difficulty. To address this, we propose TDVB, a training-free dynamic verification budget controller built upon a dual channel architecture. Specifically, a feed-forward Shape module constructs a position-wise acceptance prior from draft confidence margins to capture local spatial uncertainty, while a feedback Scale module estimates global prompt difficulty through an exponential moving average of observed acceptance lengths across rounds. The controller integrates these signals to adapt verification budgets online, effectively converting extra candidate capacity into accepted depth without offline calibration or additional training, while strictly preserving the target model’s output distribution. Extensive experiments show that TDVB achieves an average acceptance length of up to 10.18, improves end-to-end speedup by 14%–44% over strong static baselines, and reaches up to 7.40× over autoregressive decoding. Performance profiling reveals that these efficiency gains stem primarily from eliminating redundant verification rounds rather than merely reducing per-round execution overheads. Furthermore, the performance advantages scale robustly with model size, demonstrating that adaptive verification budgeting is essential to translating faster draft generation into sustained end-to-end acceleration and unlocking the true scaling potential of parallel speculative decoding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.