acceptodds
Under review as a conference paper at ICLR 2027

PARD-2: Target-Aligned Parallel Draft Model for Dual-Mode Speculative Decoding

Abstract

Speculative decoding accelerates Large Language Model (LLM) inference by using a draft model to propose candidate tokens that are verified in parallel by the target model. However, existing draft model training objectives are not aligned with the inference-time goal of maximizing consecutive token acceptance. In this paper, we build upon PARD to propose PARD-2, a dual-mode speculative decoding framework. First, we introduce Confidence-Adaptive Token (CAT) optimization, which reweights each token by the cumulative target-model confidence over its prefix, shifting the training focus from per-token accuracy to the acceptance length. Second, through stochastic gating of target hidden features, PARD-2 enables a single draft model to support both target-dependent and target-independent modes, serving an entire model family without retraining. Experiments on Llama3 and Qwen3 show that PARD-2 achieves up to 6.94 lossless acceleration in target-dependent mode and up to 5.71 in target-independent mode.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.