Teacher as Reference: Learning Real-Data Corrections for Few-Step Distillation
Abstract
Few-step distillation makes text-to-image generation efficient, but teacher matching alone does not specify how real observations should correct student predictions. We introduce Reference-Informed Non-adversarial Few-step Distillation (RINFD), which injects real-data information through two distinct training signals. Teacher-to-Data correction (T2D) learns teacher-referenced displacements from reconstruction–observation pairs on noisy bridge queries, then supplies prediction-level corrections at re-noised student outputs. Frequency-sensitive sliced Wasserstein distillation (FD-SWD) provides distribution-level supervision by aligning local, multi-scale statistics of predicted clean endpoints with real latents. Both signals act on the same clean prediction alongside teacher-guided distillation, without a discriminator. Controlled SDXL experiments support their complementarity: the default combination exceeds all tested FD-SWD-only weights on PickScore and ImageReward, and bridge sampling improves on reconstruction-only correction training. Weight sweeps reveal a balance between preference, color fidelity, and composition, with no single setting leading every metric. Across SDXL, SD3.5 Large, and FLUX.1-dev, the complete framework achieves the best average rank on two backbones and the second-best on the third among the evaluated four-step students, with the highest GenEval on both flow backbones.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.