Diffusion Drafters Can Predict Beyond Their Training Blocks
Abstract
Diffusion draft models predict fixed-length blocks of candidate tokens for parallel verification by a target model, and extending these blocks offers an opportunity to accept more tokens per round. However, naive extension can disrupt original predictions through bidirectional attention, mixing gains from added predictions with losses at the original positions. To assess the benefit of added predictions, we study whether frozen draft models can use their beyond-block predictive capability while preserving accuracy at the original positions. We introduce Native-Preserving Extrapolation (NPE), which prevents original positions from attending to added positions while allowing added positions to use representations from the original block. Across four frozen DSpark and DFlash models, NPE preserves original-position accuracy and provides useful additional predictions that increase the number of tokens accepted per round. With the original predictions preserved, we further analyze what supports or disrupts the added predictions, when these predictions accelerate generation, and what further training adds to prediction quality. Our findings show that a draft model’s useful prediction range can extend beyond its training block without sacrificing accuracy at the original positions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.