acceptodds
Under review as a conference paper at ICLR 2027

FVH-Bench: Evaluating Image Generator Safety Beyond Overtly Harmful Imagery

Abstract

Modern image generators can produce complex visual artifacts, including diagrams, interfaces, and records that combine text, graphics, and layout to convey detailed information. They may refuse requests for overtly harmful imagery yet still generate artifacts that guide dangerous actions or support deception. We call the risk arising from what such images enable, rather than only what they depict, functional visual harm (FVH). Existing safety benchmarks focus largely on overtly harmful content, such as pornography and graphic violence, leaving image generators' safety alignment against FVH insufficiently evaluated. To address this gap, we introduce FVH-Bench, comprising 669 prompts across five harmful visual functions and ten risk domains. The 11 evaluated image generators rarely refuse these requests and frequently generate artifacts that fulfill the targeted harmful functions. High scores for target harm realization and harmful artifact usability further indicate that these outputs can support the intended harmful purposes. A further analysis finds that existing safety detectors also provide limited coverage of FVH, indicating gaps in their ability to recognize harmful visual functions. Together, these findings show that evaluating image-generator safety alignment requires considering not only what images depict, but also what they may enable.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.