Learning from Reflection: On-Policy Self-Distillation for Text-to-Image Generation
Abstract
Unified multimodal models can diagnose errors in their own generations, providing image-level feedback for text-to-image post-training. However, existing methods typically use this feedback as sparse rewards to optimize the model, making it difficult to assign credit to individual generation states across timesteps. We propose Text-to-Image On-Policy Self-Distillation (T2I-OPSD), which turns textual reflections into dense state-level supervision, providing explicit velocity correction targets to improve prompt alignment. Specifically, the current model acts as both student and self-teacher, with the latter additionally conditioned on the student's generated image and textual reflection. By matching their velocity predictions at shared student-induced states, we turn image-specific feedback into dense directional supervision that shapes the student's velocity field for subsequent prompt-only generation. Furthermore, to stabilize supervision as the student evolves, we use reference-centered residual distillation to interpolate self-teacher predictions with those of a frozen prompt-only reference. We evaluate T2I-OPSD on multiple text-to-image benchmarks, achieving leading performance in compositional generation without additional sampling steps or inference memory.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.