Who Constructs the Harmful Event? Event Delegation in Conversational Image Generation
Abstract
Conversational image-generation systems develop underspecified requests into concrete scenes, helping decide what the resulting image depicts. However, existing jailbreaks typically require the attacker to specify a harmful event, sometimes leaving the target system to elaborate on it. Instead, in this work, we investigate whether an attacker can exploit this capability while leaving the concrete harmful event unspecified. We introduce Ideate-and-Concretize (I-CON), which supplies a harm-sparse scene (a scene without harm details, such as action) and a coarse category cue, then uses three fixed text turns before a single image request. The procedure requires no external attacker model, adaptive search, or repeated image attempts, requiring no access beyond the ordinary chat interface. Despite leaving the concrete harmful event unspecified, I-CON achieves a higher attack success rate than previous attack methods, and the gap widens when the baselines are restricted to the same harm-sparse inputs. Tracing event provenance in successful cases, with human validation, we estimate that 99.35% of the realized harmful events are introduced by the target and not entailed by the input. We also find that scene descriptions violating the safe image policy can reach the user even when the final image is blocked. These results motivate evaluating both the final image and the process through which its scene is constructed.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.