acceptodds
Under review as a conference paper at ICLR 2027

Revisiting the Score Difference in Distribution Matching Distillation

Abstract

Distribution Matching Distillation (DMD) trains a few-step student generator from the difference between a frozen teacher's score and an online critic's estimate of the student's score. Its quality hinges on how well the critic estimates that score, yet the critic's design is largely left to defaults. We revisit two choices—what the critic is told, and how the critic and teacher are sized—each improving few-step quality while leaving the student, and thus inference, unchanged. First, the critic sees the score-evaluation timestep but not the source noise level of each student prediction, so it cannot distinguish an expected early-step blur from a genuine error and fits a score averaged across steps. Conditioning it on the source level makes it step-aware in a few lines of code and adds no training or inference cost. Second, the student, teacher, and critic need not share one size, but this asymmetry has a limit—the critic must scale with the teacher. An undersized critic loses quality and destabilizes training; with the critic small, even a larger teacher hurts rather than helps. The critic should therefore match the teacher's size, while only the student stays small, keeping inference cheap. We validate both changes on class-conditional ImageNet, SDXL text-to-image, and WAN text-to-video in both bidirectional and autoregressive regimes, spanning -prediction and flow-matching models. In every setting, they consistently improve few-step distillation quality.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.