CAMP-IE: Understanding Safety Risks in Image Editing Models
Abstract
Multimodal instruction following image editing models are increasingly used in visual creation workflows and multimodal systems. These models combine a multimodal Transformer that interprets a source image and a natural language edit instruction with a DiT generator that performs the requested change. This flexible interface creates a safety risk because malicious intent can be distributed across the source image and an apparently benign instruction. Recent benchmarks expose this conditional failure mode. Its internal mechanisms remain unclear. We investigate the conditioning process that occurs before generation. A matched factorial design crosses sources and edits and isolates attention heads that respond selectively to harmful combinations. We call them Compositional Harmfulness Heads. We introduce CAMP-IE, which learns a low dimensional Contrastive Interaction Subspace and detects compositional risk from one conditioning pass. On IESBench, CAMP-IE yields a 48.8% bypass rate and a 42.6% rate of successful high risk edits (HRR), and retains 94.2% benign edit utility. The advantage persists for source identities and edit templates excluded from training and when the complete pipeline is refitted independently on three additional architectures. These results provide evidence that the interaction between source and edit is a recurring safety signal and that relational monitoring can operate before generation in this class of Transformer and DiT models. redWarning: This paper contains offensive images created by large image editing models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.