Domain-Edit: A Benchmark for Domain-Unbiased Evaluation of Generative Models
Abstract
As text-to-image models advance, many of the prominent compositional generation benchmarks are saturating. Yet, modern generation models have been shown to be limited in their generation diversity. This makes it difficult to assert whether the generative models have developed general abilities, or their abilities have become specialized. We observe that these models are also biased in domain (like style, layout, object properties), a phenomenon we dub generator-domain bias. Existing compositional benchmarks do not control for generator-domain bias, so models are free to fall back on their preferred domain. We argue that editing tasks serve an appropriate medium for studying this: unlike pure generation, image editing requires the model to compose within a source image it did not choose. To this end, we introduce generator-domain bias, a benchmark that pairs compositional editing instructions with source images generated across diverse styles. We evaluate eight open-source models on three subsets of increasing visual diversity, spanning plain textures and five styles. We find that every model falls below its own generation accuracy once the domain is fixed, that performance varies substantially with the source style, and that models often regenerate the background rather than edit it. Based on these findings, we argue that input-domain diversity should be a core dimension of compositional evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.