Structural Illumination under Control–Realization Mismatch for Compositional Text-to-Image Generation
Abstract
Compositional text-to-image generation is typically optimized to produce a faithful image, yet the same compositional constraint can admit many valid realizations with different spatial configurations. We study the complementary problem of : given a fixed generation budget, how can we uncover diverse realizations that satisfy the same compositional constraint? We formalize this objective as , which seeks to cover distinct structural niches using constraint-satisfying generated images. A key challenge is : inference-time search selects among candidate control configurations that specify intended structure, but stochastic generation may realize a different structure in the resulting image. To address this mismatch, we introduce Archive-Guided Structural Illumination (AGSI), which uses cheap structural proxies derived from these controls to guide search. Only constraint-satisfying generated images can expand structural coverage, and their realized structure determines the niche they occupy. On GenEval spatial relations with GLIGEN, AGSI achieves higher realized structural coverage than strong search baselines under matched generation budgets. The same principle transfers to exact counting and T2I-CompBench++ prompts, and remains effective when replacing GLIGEN with InstanceDiffusion. A matched ablation further shows that tracking discovered niches from realized outputs rather than proxy predictions improves coverage on both tasks. Notably, AGSI achieves broader coverage than the strongest baseline with a similar or lower fraction of feasible generations, indicating that its gains come from using feasible generations more effectively rather than simply producing more of them.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.