acceptodds
Under review as a conference paper at ICLR 2027

FairFlow: Demystifying and Mitigating Stereotype Bias in Text-to-Image Diffusion Transformers

Abstract

Multimodal diffusion transformers (MM-DiTs) underpin modern text-to-image systems, yet can reproduce demographic bias under neutral prompts. Existing debiasing methods were largely designed for U-Net backbones and offer limited guidance on where semantic attributes are controlled in joint transformer architectures. In this work, we trace stereotype bias through MM-DiTs and find that its causal influence is concentrated in a sparse set of semantic control layers. These layers exhibit strong visual and cross-modal signals, as revealed by our attention analysis. Crucially, our investigation into text-state trajectories uncovers a pivotal bias reinforcement loop, demonstrating how repeated text-image updates systematically amplify initial bias priors during the generation process. Based on these observations, we propose FairFlow, a parameter-preserving method that learns reusable attribute directions and injects them into visual features at the identified control layers during early denoising. Experiments on FLUX.1 and Stable Diffusion 3 across multiple demographic settings show that FairFlow improves the fairness–fidelity tradeoff and remains effective on complex scene prompts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.