The Prompt Effect: How Conditioning Strength Shapes Stereotype Differentiation in Text-to-Image Diffusion Models
Abstract
Bias audits of text-to-image diffusion models typically evaluate generations produced under explicit semantic prompts, even though measured demographic differentiation may itself depend on how strongly those prompts condition generation. We study this by comparing an operational weak-conditioning baseline, empty-prompt generation at classifier-free guidance (CFG) scale 1, against neutral prompted generation across six gender–occupation and social contexts. On Stable Diffusion v1.5, prompted generations show consistently greater per-image stereotype differentiation than the weak-conditioning baseline (LBAS < 1 in all six contexts), robust to relevance filtering and normalization. We then intervene on conditioning strength, holding prompt and seed fixed. Increasing CFG from 1.0 to 7.5 increases |skew| by 0.00747 on average across 36 matched context–seed pairs (two-sided Wilcoxon p = 0.00136), and a context fixed-effects regression estimates a positive CFG effect (β = 0.001079, p = 0.00121). However, CFG also increases semantic relevance, and the CFG–skew association attenuates to approximately zero after adjusting for relevance, suggesting the amplification is closely coupled to stronger semantic alignment rather than a relevance-independent direct effect. The directional pattern also replicates on SDXL and under OpenCLIP rescoring. These results characterize stereotype differentiation as a property of the conditioning pipeline, not just the underlying model, and motivate fairness audits that evaluate behavior across conditioning strength rather than at one prompted operating point.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.