acceptodds
Under review as a conference paper at ICLR 2027

LatentCompass: T2I Diffusion Steering via Orthogonal Attribute Spaces for Debiasing, Concept Erasure, and Red Teaming

Abstract

Text-to-image (T2I) diffusion models suffer from biased results stemming from entangled generative priors and a lack of accurate control over outputs. Current mitigation attempts rely on imprecise, adversarially-vulnerable prompt and text embedding interventions, or they require prohibitive and invasive fine-tuning. Further, text-based methods can only control descriptive attributes, i.e., what an image depicts, but not evaluative attributes, i.e., how it is perceived by an external judge. We propose LatentCompass, an exemplar-based approach that enables disentangled and controllable generation for both descriptive and evaluative concepts in a training-free manner. LatentCompass steers the generative trajectory by (a) constructing a nonlinear, low-dimensional, and orthogonal attribute space via a closed-form solution that explicitly isolates desired concepts, (b) defining a target- class intervention in this space, and (c) reflecting the corresponding intervention directly in the T2I latent space. Extensive evaluations demonstrate that LatentCompass on average (i) mitigates generative stereotypes in medical professions by 100%, (ii) reduces unsafe concept generation to 1.3%, (iii) enhances aesthetic score by 27%, (iv) boosts red-teaming success rates against Deepfake detectors by 40%, and (v) enables high-fidelity attribute editing without leakage.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.