ProOPD: Process-Guided On-Policy Distillation for Iterative Text-to-Image Generation
Abstract
Unified multimodal models (UMMs) can improve text-to-image generation through an iterative process that interleaves textual reasoning, image synthesis, reflection, and refinement, but training such policies faces two challenges: heterogeneous task objectives and the limited guidance that final-image rewards provide for intermediate decisions. We introduce ProOPD, a process-guided on-policy distillation framework for iterative text-to-image generation. A shared student produces its own text and image trajectories, and frozen stage-specific teachers provide dense supervision at the resulting token prefixes and visual sampling states. The supervision covers initial reasoning and generation as well as subsequent reflection and refinement, following the states encountered by the student as its policy changes. We combine textual distribution matching and visual velocity matching with image-level outcome optimization to train the generation process toward improved image quality. Experiments on four challenging text-to-image benchmarks (OneIG-EN, LongText-EN, WISE, and GenEval++) assess generation performance and the effectiveness of iterative refinement, with ablations examining the contributions of process supervision and outcome optimization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.