Bias Before Pixels: A Surrogate Framework for Auditing Demographic Bias in Text-to-Image Models
Abstract
Text-to-image (T2I) models reproduce demographic biases from their training data, but auditing this bias traditionally requires generating and classifying large image batches for every evaluated prompt. We train a surrogate that predicts the demographic distribution of a model's generated images directly from the prompt, so new prompts can be audited without generating any images. The surrogate can reuse the target model's own text encoder, or use an external encoder that requires no access to the model's weights, making it also suitable for API-only systems. We construct a dataset of long and short prompts paired with measured gender and race distributions for three T2I models (FLUX.2-klein, SDXL Turbo, Z-Image Turbo), on which our surrogate improves gender bias-detection accuracy by 11–25% over the strongest text-based or diffusion-internal baseline, with consistent gains on race. Since each audit takes a single forward pass, we can compare bias across models and find concepts and words whose demographic skew differs sharply between them. This efficiency also lets us test all 9,396 object–activity combinations, which shows that activities shape gender bias more than objects. Finally, the surrogate improves an existing debiasing pipeline by predicting failed interventions before image generation, raising gender fairness from 0.62 to 0.74 without reducing image quality. Surrogates thus make it practical to audit bias at scale, before generating a single image and even without access to model weights.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.