acceptodds
Under review as a conference paper at ICLR 2027

PromptEcho: Annotation-Free Reward from Vision-Language Models for Text-to-Image Reinforcement Learning

Abstract

Reinforcement learning can improve prompt following in text-to-image (T2I) models,but obtaining high-quality reward signals remains a key challenge. Existing rewards either provide only coarse-grained global alignment signals or rely on costly human preference annotations and additional reward-model training. We introduce PromptEcho, an annotation-free reward that uses a frozen vision-language model (VLM) to compute the image-conditioned, teacher-forced next-token likelihood of the original prompt. This directly converts the VLM’s pretraining objective into a continuous token-level reward without training a separate reward model. For evaluation, we develop DenseAlignBench, a benchmark of concept-rich dense captions designed to rigorously test fine-grained prompt following. Relative to the corresponding original baselines, PromptEcho achieves net pairwise advantages of +26.8pp and +16.3pp on DenseAlignBench for Z-Image and QwenImage-2512, respectively, and yields consistent improvements on public benchmarks including GenEval, DPG-Bench, and TIIFBench without any task-specific training. Under the same Z-Image RL configuration, PromptEcho also substantially outperforms CLIP score, ImageReward-v1, and HPSv3, with a particularly pronounced advantage on long, dense prompts. Ablations with VLMs of different sizes further show that PromptEcho can directly benefit from advances in open-source VLMs. Code, checkpoints, and DenseAlignBench will be released through an anonymous project page.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.