acceptodds
Under review as a conference paper at ICLR 2027

Who Constructs the Harmful Event? Event Delegation in Conversational Image Generation

Abstract

Conversational image-generation systems develop underspecified requests into concrete scenes, helping decide what the resulting image depicts. However, existing jailbreaks typically require the attacker to specify a harmful event, sometimes leaving the target system to elaborate on it. Instead, in this work, we investigate whether an attacker can exploit this capability while leaving the concrete harmful event unspecified. We introduce Ideate-and-Concretize (I-CON), which supplies a harm-sparse scene (a scene without harm details, such as action) and a coarse category cue, then uses three fixed text turns before a single image request. The procedure requires no external attacker model, adaptive search, or repeated image attempts, requiring no access beyond the ordinary chat interface. Despite leaving the concrete harmful event unspecified, I-CON achieves a higher attack success rate than previous attack methods, and the gap widens when the baselines are restricted to the same harm-sparse inputs. Tracing event provenance in successful cases, with human validation, we estimate that 99.35% of the realized harmful events are introduced by the target and not entailed by the input. We also find that scene descriptions violating the safe image policy can reach the user even when the final image is blocked. These results motivate evaluating both the final image and the process through which its scene is constructed.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.