When Mismatches Overlap: Concurrent Verification for Speculative Decoding
Abstract
Speculative decoding (SD) accelerates large language model inference by verifying multiple draft tokens in a single target-model pass. Strict verification stops at the first mismatch and discards all remaining drafts, even though the target model has already computed predictions for later positions. Deferred verification addresses this waste by treating subsequent target outputs as evidence: instead of rejecting a mismatch immediately, the verifier continues and resolves it based on what follows. However, a later mismatch can arise before an earlier one is resolved, leaving multiple verification decisions pending simultaneously. Existing methods focus on the earliest pending mismatch: a later mismatch triggers rejection instead of becoming an independent verification decision, limiting draft reuse. We introduce PEND, a concurrent verifier that maintains a separate state for each pending mismatch while sharing subsequent target predictions across them. PEND adapts evidence requirements to target-relative support, weights subsequent observations by target confidence, and commits a draft segment only after all pending mismatches resolve. Using only cached target outputs from the current verification pass, PEND improves throughput by 46.6-70.3% (57.2% on average across the evaluated settings) relative to vanilla SD across two model families and four benchmarks, while keeping task quality close to strict verification and achieving better quality-efficiency trade-offs than recent deferred-verification baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.