CSD: Confidence- and Stability-Guided Speculative Decoding in Large Language Models
Abstract
Speculative decoding has emerged as a promising technique to accelerate large language model inference by leveraging a lightweight draft model to propose candidate tokens for parallel verification by a target model. Although recent methods have explored adaptive draft-length control, efficiently estimating when the draft model can safely speculate more aggressively remains challenging. In this paper, we propose Confidence- and Stability-guided speculative Decoding (CSD), a training-free method that relies only on draft-model runtime signals, including token-level prediction confidence and recent local stability. CSD adopts a lightweight two-level framework to dynamically adjust the drafting boundary during decoding. Experiments across multiple model pairs and datasets show that CSD can be integrated into both standard and parallel speculative decoding without modifying the target verification rule and achieves a – speedup over autoregressive decoding. CSD achieves competitive or better throughput than strong fixed-length and parallel speculative decoding baselines across the evaluated settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.