acceptodds
Under review as a conference paper at ICLR 2027

BridgePrompt: Bridging the Implicit Meaning Gap in Text-to-Image Generation

Abstract

Text-to-image (T2I) models can render explicit prompts well, but often fail when the desired image depends on meaning that is only implicit in the user's request, such as figurative meaning, entity knowledge, physical consequences, or scientific relations. We call this failure mode the implicit visual meaning gap. We view it as a conditioning-interface problem: the prompt wording may not make explicit the visual semantics that the generator must realize, even when those semantics are recoverable from the request. This motivates a prompt-side objective of making the intended visual meaning explicit before generation. We propose BridgePrompt, which keeps both the language prompter and T2I generator frozen and optimizes only a reusable natural-language instruction using feedback on generated images from a frozen vision-language evaluator. Candidate instructions are scored by how well their generated images match the original prompt, rather than by the rewritten text itself, while inference requires only one refinement and one generation. Conditioning the generator on benchmark-provided intended-meaning annotations raises the score to , supporting the view that many failures can be repaired by making missing semantics explicit before generation. With a frozen 1.5B prompter, BridgePrompt improves SD3.5 Medium on T2I-ReasonBench from to , outperforms zero-shot GPT-4o refinement, and transfers without modification to other reasoning benchmarks and T2I models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.