acceptodds
Under review as a conference paper at ICLR 2027

Rescuing Regrettable Rejections: Relaxed Acceptance for Speculative Decoding

Abstract

Speculative decoding has become one of the most effective techniques for accelerating inference in large language models. Its defining property is that it is lossless: the output distribution is identical to that of the target model, so the speedup comes at no cost in output quality. In speculative decoding, a small draft model proposes several tokens, the larger target model verifies them in parallel, and decoding resumes from the first proposal that fails verification. Every draft token after the rejected one is discarded, and the more tokens are discarded, the smaller the speedup becomes. We examine these rejections and find that many of them are regrettable. A rejection is regrettable when accepting the draft token would not have affected the quality of the output. We propose a method that rescues such rejections without training any additional model or adapter. It targets the three types of regrettable rejection that we identify, surface form, preference, and order, with a whitelist of interchangeable tokens, mutual confirmation between the two models, and order-swap tables, respectively. Experiments on Qwen3-14B and DeepSeek-R1-Distill-Qwen-32B, each drafted by a smaller 4-bit quantized model of the same family, show that the method rescues 30–36% of the rejections of lossless speculative decoding at temperature 0 and 14–15% at temperature 0.6, which raises throughput by a further 8–13% and 3.5–6.5%, respectively, and lifts the speedup over the target alone to 2.1–2.4×, with no forward pass added and both models left untouched. Mean accuracy stays within 0.4 points of lossless decoding; the one clear loss, on a single benchmark, is traced to one rule.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.