acceptodds
Under review as a conference paper at ICLR 2027

From Prompt Alignment to Intent Recovery: Rethinking Text-to-Image Evaluation

Abstract

Text-to-image models have made substantial progress in faithfully rendering creators' prompts. However, ambiguous visual cues may lead observers to misinterpret the creator's intent from images alone, potentially triggering public alarm, reputational damage, and economic losses. Existing prompt-conditioned metrics primarily evaluate creator-side prompt alignment but do not directly measure observer-side intent recovery. Hence, we introduce **Prompt-Agnostic Intent Recovery** (PAIR), a protocol for evaluating intent recovery from generated images alone. A prompt-agnostic observer produces an open-ended image description and a separate scorer evaluates the description against a pre-specified intent rubric. We construct a benchmark of 300 intents and 1,800 images from two generators. Across the benchmark, prompt alignment and intent recovery induce different candidate rankings. When selecting among six images per intent from two generators, PAIR-based selection achieves 9-17% higher recovery on held-out evaluations than CLIP, VQAScore, and a target-aware VLM selector. In diagnostic selector-disagreement cases, participants without prompt access achieved 85.9% two-choice accuracy for PAIR-selected images versus 39.4% for alignment-selected images. Finally, PAIR-guided refinement improves the mean final PAIR score by 23% over repeated uncalibrated generation. These results identify observer-side intent recovery as an evaluation target not captured by prompt-conditioned alignment.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.