BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding
Abstract
Speculative decoding accelerates inference by using a draft model to generate candidate tokens in parallel, which are then verified by the target model, enabling lossless acceleration. Recently, diffusion-based speculative decoding further improves parallelism by generating multiple tokens per forward pass via block-level diffusion, achieving state-of-the-art (SOTA) performance. However, existing methods adopt a fixed inference block size and assume a uniform optimal decoding strategy across all inputs. In this paper, we show that this assumption is suboptimal, as the optimal block size varies across samples and plays a critical role in speculative decoding performance. Moreover, these values exhibit a clear local structure, concentrating around the training block size, which reduces the problem to a low-dimensional and structured decision space. Based on these insights, we propose BlockPilot, a sample-adaptive block size selection method. Specifically, we formulate block size selection as a policy learning problem and propose an instance-adaptive decision mechanism that predicts the optimal block size based on the representation of the prefill stage. The prediction is performed only once after prefill. Extensive experiments demonstrate that BlockPilot consistently improves efficiency, achieving an acceptance length of 5.92 and a 4.20 speedup on Qwen3-4B under temperature .
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.