A2I-Bench: Measuring the Value of Clarification in Interactive Text-to-Image Generation
Abstract
Interactive text-to-image systems ask for clarification when a user prompt admits multiple plausible interpretations. Yet ambiguity alone does not determine whether clarification is useful: the same user intent may already match one generator's behavior under the original prompt while conflicting with another's. To formalize this issue, we introduce A2I-Bench, a new benchmark of 527 ambiguous text-to-image prompts for measuring the value of clarification (VoC). Each case contains two linguistically plausible, visually distinguishable user intents together with intent-specific visual contracts, enabling matched counterfactual evaluation while holding the prompt, generation system, and image budget fixed. We operationalize clarification as revealing the user's intended interpretation. For each generation system, we estimate a Direct-based default interpretation from independent direct generations and measure the gain from intent revelation relative to a frozen no-question rewrite. Across four generation systems, revealing the intended interpretation improves contract satisfaction by 0.158 when the user's intent conflicts with the Direct-based default, a 0.092 larger gain than when they are aligned. These results show that clarification value depends on whether the no-question generation behavior already realizes the user's intent, rather than on prompt ambiguity alone. Because generator defaults and clarification gains are estimated separately for each system, A2I-Bench provides a reusable protocol for evaluating clarification value in new text-to-image generators. We will release the benchmark and code.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.