PE-OPSD: Internalizing Prompt Enhancement into Flow-matching Models via On-Policy Self-Distillation
Abstract
Text-to-image users often provide short and underspecified prompts, while generative models rely on detailed textual conditions for reliable instruction following. Existing systems mitigate this mismatch with Prompt Enhancers (PEs), which rewrite raw prompts into richer descriptions at inference time, but this introduces extra cost and leaves prompt understanding external to the generator. In this work, we reinterpret prompt enhancement as privileged information and ask whether its benefit can be internalized into the generator itself. We propose Prompt-Enhanced On-Policy Self-Distillation () for text-to-image flow-matching models. During training, a student conditioned only on the raw prompt follows its own generation trajectory, while an enhanced-prompt teacher provides vector-field supervision at the states visited by the student. This on-policy formulation transfers enhanced-prompt generation behavior into the raw-prompt student without using enhanced prompts or offline-generated samples at inference time. After training, both the PE and teacher are removed. Experiments across multiple models and benchmarks show that improves prompt fidelity and visual quality while reducing inference-time prompt-rewriting overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.