acceptodds
Under review as a conference paper at ICLR 2027

Target Correction for Diffusion Sampling and Speculative Decoding

Abstract

During LLM inference, control algorithms decide which tokens to resample for diffusion models and which request tokens to verify for speculative decoding. For the sake of efficiency and simplicity, these algorithms do not attempt long-horizon planning to hit a specified target—fidelity for diffusion sampling, or goodput for speculative decoding. This paper presents *target correction*, a technique for noninvasively imbuing simple control algorithms with target awareness. The approach (1) interprets the algorithm as myopically playing a betting game in which wealth multiplies according to random outcomes, (2) devises a target-aware utility function of wealth, (3) derives a target-aware value function from a continuous-time approximation of the game, and (4) uses that value function to modify the original algorithm. Our correction of the Entropy-Bound diffusion sampler in DiffusionGemma achieves 4% lower NFE at comparable or better fidelity. To illustrate how surgical target correction can be, we improve DSpark's goodput robustness under skewed workloads by up to 10% with just a single-line change.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.