Seen but Unnamed: How Caption Coverage Affects Composition in Diffusion Models
Abstract
Compositional generation asks a model to satisfy a prompt that names several objects at once, each with its attributes and relation to the scene, and diffusion models occasionally and unpredictably fail at such a task. Two kinds of accounts attempt to explain why, in the hopes of arriving at a broader claim of where compositionality originates. Architectural claims attribute composition to a property of the denoiser such as locality, under which each output pixel depends only on nearby pixels. Data-centric claims attribute it to the training data, most often to image coverage, the attribute combinations in the training images. However, a diffusion model also learns from its captions, and human captions leave most of the objects in an image unnamed. We introduce and vary caption coverage, the fraction of multi-object training images whose caption names at least two of their objects, while holding other experimental settings constant. We find that diffusion models follow a prompt that names several objects only when enough of their captions name several objects together, while image coverage mainly governs transfer to unseen color and shape pairs. We restore composition by summing a failing model's single-object predictions, and conclude that the value of captions lies in teaching a model to read several objects from one prompt. Finally, across architectural ablations, we find that cross-attention and sampler design set the caption coverage a model needs, and locality supports prompts with more objects without predicting which models trained on the same data succeed. In contrast, anyone can count caption coverage before training and raise it by rewriting captions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.