acceptodds
Under review as a conference paper at ICLR 2027

TIDIER: Text-to-Image Alignment via Knowledge-Guided Hate Detection, Interpretation, and Repair

Abstract

Text-to-Image diffusion models are widely deployed but remain unsafe, generating hateful stereotypes, demographic bias, and dehumanizing compositions under both adversarial and benign prompts. Existing safety methods rely on endpoint filtering or global trajectory steering over unstructured signals, leading to over- censorship, poor generalization, and failure on compositional harms. We instead introduce the Semantic Safety Margin (SSM), a structured similarity constraint over a Multimodal Knowledge Graph (MMKG) encoding harmful concepts, benign counterparts, and contextual relations. Building on this, we propose TIDIER, a training-free framework with dual-space alignment: prompt-space repair and trajectory-space intervention. Open-World Concept Inference (OWCI) detects novel harmful semantics and proposes them for human-reviewed graph expansion. Across four benchmarks (Detonate, I2P, T2I-Safety, T2I-RiskyPrompt), TIDIER reduces prompt-level toxicity by 65% and image toxicity by up to 64% while achieving a 9–10% relative gain in CLIP score. The code and MMKG are available at https://anonymous.4open.science/r/TIDIER-ICLR-148D/ .

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.