Just as Fair: Balancing Diffusion Trajectory for Fair text-to-Image Generation
Abstract
Diffusion models exhibit persistent demographic biases, systematically favoring some attributes over others in text-to-image generation. While prior work largely address bias in prompts or representations, we ask a different question: does the denoising process itself explore demographic attributes unevenly? We find that it does, biased generation is closely associated with an imbalance in the exploration of denoising trajectories, which reflects statistical biases in the training data and persists throughout sampling. This motivates a trajectory-level view of fairness, where fair generation is treated as balanced exploration of demographic attribute modes, rather than solely as a property of the final outputs. Based on this view, we develop a trajectory-level debiasing framework for fair text-to-image generation. We first introduce trajectory bias, together with an observation mechanism based on Tweedie's formula, which estimates the semantic outcome of intermediate noisy latents and makes attribute preferences measurable throughout denoising. Building on this signal, we dynamically regulate the denoising trajectories with a balanced allocation budget, steering the sampling process toward fair attribute coverage. A key advantage of our approach is its architecture-agnostic nature, which intervenes directly in sampling dynamics and overcomes a core bottleneck of existing debiasing methods. Across four representative diffusion architectures, including U-Net, DiT, and rectified flow models, our method consistently improves demographic balance while preserving generation quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.