acceptodds
Under review as a conference paper at ICLR 2027

SpecPE: Reinforcing Structured Reasoning for Prompt Enhancement

Abstract

Prompt enhancement (PE) for reasoning-intensive text-to-image generation requires translating implicit user intent into explicit scene semantics. However, image-level rewards alone provide limited feedback on whether errors arise from scene interpretation or image synthesis. We introduce SpecPE, a modular PE framework that makes the model's declared scene interpretation directly evaluable during reinforcement learning. Its policy generates a rationale, a structured scene specification, and an enhanced prompt, which alone conditions a fixed image generator. A structure-aware evaluator parses the specification's entities, attribute bindings, and relations, assesses their alignment with the user request, and checks consistency with the enhanced prompt. To this end, we construct SpecPE-CoT, a dataset of 32,599 synthetic structured demonstrations, for supervised fine-tuning, followed by group relative policy optimization with a completion-level reward combining semantic and image feedback. Across three combinations of PE backbones and image generators, SpecPE improves overall T2I-ReasonBench accuracy by up to 6.87% over the corresponding PE baselines. On the WISE benchmark, our proposed SpecPE reaches 0.65, exceeding the strongest evaluated baselines by 0.07 and 0.08. Sequential reward ablations show gains from adding semantic feedback and further gains from image feedback.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.