Draw What LLMs Intend: Prompt Alignment Is Not Enough for Text-to-Image Generation
Abstract
In Text-to-Image (T2I) generation, which image should we generate among multiple candidates that all satisfy the same prompt? T2I generation is no longer a prompt-to-image black box, but a pipeline in which an upstream model resolves semantics before a generator renders the image. This shift matters because prompts are structurally under-specified and permit multiple images that all preserve the same prompt anchors. Yet prior work primarily measures visual quality or prompt-image alignment, leaving unaddressed whether the generator faithfully realizes the semantic decisions formed by its upstream model. We argue that, among prompt-satisfying images, the visual commitment formed by the upstream model should be carried through to the final image; we call this problem Latent Intent Transfer (LIT). In this paper, we study LIT as a new axis for evaluating and improving T2I generation systems. We formulate LIT as the second-stage faithfulness problem of a modern T2I pipeline: whether the generator realizes the upstream model's resolved visual commitments. For evaluation, we introduce two metrics, intent win rate and open-slot commitment accuracy, along with DaLI-Bench, a controlled open-slot benchmark where multiple prompt-valid images remain possible. Guided by the four desired properties, we propose DaLI, a plug-and-play method for improving existing T2I pipelines. Extensive experiments show that DaLI consistently improves LIT while preserving prompt-anchor fidelity and visual quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.