Compositional Collapse in Diffusion Models: Scaling Laws of Local Fidelity
Abstract
That diffusion models struggle with crowded scenes is known; how they fail, including the functional form of the failure, its onset, and the extent to which it is a measurement artifact, has never been systematically measured. We introduce a controlled framework that varies the requested entity count from 1 to 64 while keeping scene, lighting, and activity fixed, measuring local fidelity (entity recall, facial validity, identity duplication) alongside global quality (CLIPScore and per-density-bin KID). Three results emerge. (1) Collapse is exponential, not gradual: across three open models (SDXL, SD1.5, and PixArt-α), detector-corrected recall decays exponentially with entity count and is preferred over power-law fits by AIC in all models, with decay constants ranging from 0.019 to 0.034. This yields a characteristic capacity scale, corresponding to a recall half-life of 20 to 36 entities, rather than scale-free degradation. The same collapse pattern reproduces for cars instead of people, with rate constants unchanged in two of the three models, suggesting a largely class-independent bottleneck. (2) Naive measurement overstates collapse by approximately 40%: even a strong face detector misses more than one-third of ground-truth faces in dense real scenes, with recall dropping from 0.94 to 0.59 on WIDER FACE. We account for this confound through calibration on real-image controls. (3) Identity duplication has a measurable onset: it is rare below roughly 8 entities (0 to 11% of images) but appears in 75 to 84% of images at 64 entities using an ArcFace cosine similarity threshold greater than 0.4. Segmented regression estimates the onset between approximately 10 and 22 entities. Meanwhile, global metrics are either insensitive to or negatively correlated with collapse: CLIPScore remains flat or increases, while KID against real crowds improves. These findings reveal a decoupling between global realism and local structural correctness that remains invisible to standard evaluation metrics.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.