acceptodds
Under review as a conference paper at ICLR 2027

StateMix: Complementary State Sources for Few-Step Diffusion Distillation

Abstract

Few-step diffusion distillation has advanced through a variety of training objectives. Beyond the objective itself, the states on which a student is trained also shape the distillation process. We introduce noise-start states, sampled directly from a fixed Gaussian prior at intermediate target times, as a new training-state source for distribution matching. Controlled toy experiments reveal complementary learning behaviors: noise-start training strengthens responses to changed conditions, whereas rollout-start training on states produced by the student’s preceding denoising steps transmits more input-state variation. Building on this complementarity, we propose StateMix, a mixed-state training strategy that combines both sources under the same DMD2 supervision without changing inference. On text-conditioned SDXL and Qwen-Image, StateMix achieves leading human preference scores among the evaluated methods, with the mixture ratio allowing a trade-off between preference scores and similarity to the teacher’s generated distribution. We further evaluate StateMix on Qwen-Image-Edit with both text and image conditions, where public benchmarks demonstrate its effectiveness. In human GSB evaluation, the student meets all eight application-specific quality criteria relative to the teacher, with no loss under the production acceptance standard and even surpassing the teacher on most evaluation dimensions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.