Parachuting parrot: A dataset for evaluating visual compositional generalization
Abstract
Proprietary text-to-image models have become remarkably powerful at rendering complex compositional scenes. Did they overcome the longstanding challenge of compositional generalization, and if so, how did they do it? Current datasets that allow to systematically evaluate compositional generalization lack the complexity, diversity and scale required to train image generation models of remotely similar quality to answer this question. Here, we introduce Parachuting Parrot, a vision dataset and benchmark for evaluating compositional generalization consisting of over 2.5 million images with corresponding structured scene descriptions. Parachuting Parrot consists of complex scenes defined by a rich scene description language, rendered using large-scale text-to-image models and validated through a scalable vision-language model pipeline. Alongside it, we define systematic splits to evaluate compositional generalization over scenes with varying compositional complexity. We train a text-to-image latent diffusion model from scratch on Parachuting Parrot, achieving generalization to complex scene compositions with multiple entities, entity-attribute binding and relational binding in a fully controlled experimental setting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.