acceptodds
Under review as a conference paper at ICLR 2027

Dispersed Any-Order Speculative Decoding for Exact Parallel Autoregressive Image Generation

Abstract

Autoregressive image models generate one token per forward pass, which makes inference slow. Speculative decoding can speed up inference by generating multiple tokens per pass without changing the sampled distribution. The speedup depends on how much the drafter's distribution disagrees with the target's, and we show that this disagreement has two sources: 1) the drafter can be a different model than the target, and 2) it can condition on different tokens than the target since the drafter generates in parallel while the target conditions each draft on the ones before it. Recent self-speculative methods use the target model as its own drafter, which removes the first source. However, for left-to-right generation, since the drafted positions are adjacent and strongly dependent on one another, the second source of disagreement remains large. We find that disagreement decreases as distance between draft tokens increases, reducing the second source of disagreement. Based on this, we propose Dispersed Any-Order Speculative Decoding (DSD), which uses an any-order model as its own drafter and verifier along an order whose consecutive positions are far apart. DSD requires no additional training and samples from the exact target model's distribution. On three any-order image models, DSD samples the exact distribution with up to fewer forward passes than sequential decoding, fewer than the best existing self-drafting method, and fewer than speculative decoding with a trained drafter, and it matches the quality of the models' lossy sampling with – fewer steps.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.