acceptodds
Under review as a conference paper at ICLR 2027

What Goes Unsaid: Measuring Unsafe Visual Defaults in Text-to-Image Generation

Abstract

We study a systematic bias in text-to-image models: when a benign prompt leaves a safety-critical visual attribute unspecified, the model may complete it with a norm-violating configuration. We term this phenomenon default-unsafe behavior and introduce UNSAID, a source-grounded benchmark of 498 core norms and 1,144 benign prompts with executable four-way evaluation rubrics. Across nine T2I models, default-unsafe behavior is widespread: even the best-performing model violates the applicable norm in 23.1% of decisive generations, with violations spanning all four safety domains and seven visual-attribute types. Generic safety instructions provide limited benefit, while explicit specifications reduce but do not eliminate violations. Repeated sampling increases both the likelihood of obtaining a safe image and the likelihood of encountering a violation, and existing image guards miss most violations, with the best configuration missing 75.6%. Building on these findings, we develop an agentic safety harness framework that reduces the average violation rate by 15%. Together, these findings establish default-unsafe behavior as a distinct and largely unaddressed safety problem in image generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.