FVH-Bench: Evaluating Image Generator Safety Beyond Overtly Harmful Imagery
Abstract
Modern image generators can produce complex visual artifacts, including diagrams, interfaces, and records that combine text, graphics, and layout to convey detailed information. They may refuse requests for overtly harmful imagery yet still generate artifacts that guide dangerous actions or support deception. We call the risk arising from what such images enable, rather than only what they depict, functional visual harm (FVH). Existing safety benchmarks focus largely on overtly harmful content, such as pornography and graphic violence, leaving image generators' safety alignment against FVH insufficiently evaluated. To address this gap, we introduce FVH-Bench, comprising 669 prompts across five harmful visual functions and ten risk domains. The 11 evaluated image generators rarely refuse these requests and frequently generate artifacts that fulfill the targeted harmful functions. High scores for target harm realization and harmful artifact usability further indicate that these outputs can support the intended harmful purposes. A further analysis finds that existing safety detectors also provide limited coverage of FVH, indicating gaps in their ability to recognize harmful visual functions. Together, these findings show that evaluating image-generator safety alignment requires considering not only what images depict, but also what they may enable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.