acceptodds
Under review as a conference paper at ICLR 2027

Scored Consistency Distillation: Choosing Which Teacher Trajectory to Imitate

Abstract

Few-step distillation trains a student to imitate a teacher's sampling trajectory, but which trajectory the teacher produces is decided by the seed. For a single caption a teacher generates outcomes that differ widely in how faithfully they realise the prompt, and standard distillation imitates one of them at random, inheriting the teacher's compositional failures as readily as its successes. We show that choosing the trajectory to imitate, and adding a reward on the student's own clean-latent estimates, improves compositional alignment with low online overhead. For each caption the frozen teacher samples several candidates; we select the one whose mean-pooled DINOv2 patch embedding is closest to the caption's real photograph, and distil on that trajectory alone. A lightweight projector maps the student's terminal latent directly into DINO space, so the same score applies to the student's own predictions. The projector tracks the true score on the student's recent predictions, and the two are updated on separate steps, which prevents direct co-adaptation between the reward and the policy it scores. On SD3.5-Medium distilled to four steps, the method improves T2I-CompBench++ from to and GenEval2 from to , with the largest gains on the spatial categories, followed by texture and numeracy. The four-step student surpasses the ten-step teacher it was distilled from and matches the teacher at its recommended 40-step setting, at a fraction of the inference budget.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.