acceptodds
Under review as a conference paper at ICLR 2027

T-DIG: Text-Conditioned Diversity in Image Generation

Abstract

Text-to-image (T2I) platforms typically return several images per prompt so users can explore different interpretations and choose one that matches their intent. Most existing methods for diversity optimize pairwise distances in a latent embedding space, which are hard to map to interpretable attributes and don't account for user preferences. Some users may want variations in the background, whereas others may want variations in artistic styles. In this paper, we introduce T-DIG, a diversity framework for image generation and editing. T-DIG is an iterative framework comprising of multiple modules that lets users type prompts, inspect default generations, request for variations using natural language, and continue refining outputs until they're satisfied. At its core, T-DIG uses a LLM policy that produces axis-conditioned prompts or edit instructions for a frozen generator or editor. We train this policy with GRPO using a text-conditioned diversity reward and alignment penalties. Because the renderer remains frozen, T-DIG applies to open-weight, distilled, and proprietary models. We collect human-written axes to build training and validation data, and derive a test set from user interactions with \method, all of which we release as the first dataset for text-conditioned prompt diversity in image generation. T-DIG far exceeds baselines in improving text-conditioned diversity while preserving base-prompt alignment across a variety of metrics and human judgements. Our code and data is here: https://anonymous.4open.science/r/t-dig-paper

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.