Parallel-Parallel Speculative Decoding
Abstract
Speculative decoding accelerates memory-bandwidth-bound inference of large autoregressive target models using a small, fast drafter model. Target-latent-conditioned parallel drafters, such as DFlash, currently offer the best accuracy-latency tradeoff, leading to the highest reported speedups. However, latent conditioning still requires executing drafting and verification sequentially, placing drafting latency on the critical path. This limits achievable speedups and prevents effective use of emerging hardware that can support concurrent drafting and verification. We introduce parallel-parallel speculative (PaPaSpec) decoding, which runs a multi-layer latent-conditioned parallel drafter concurrently with verification, hiding drafting latency entirely. The key enabler is a prefix and depth-causal latent conditioning scheme that restricts each drafter layer to only access the target latents computed so far. Hiding drafting behind verification not only removes it from the critical latency path, but also buys the drafter additional time to refine its draft. We exploit this by looping over intermediate layers until new conditioning information arrives, which increases the acceptance length while keeping the drafter size constant. We evaluate PaPaSpec in single-batch serving across three different target models on NVIDIA H100 GPUs. Across seven benchmarks covering math, coding, and chat, PaPaSpec accelerates autoregressive generation by 3.26–5.50x on average, a 13–22% improvement over DFlash.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.