When Does an Edit Become Unsafe? Tracing Safety Emergence Along the Generative Trajectory of Image Editing Models
Abstract
Modern image editors do not produce an edit in one step: a source image and instruction are progressively transformed through a diffusion or flow trajectory. Yet existing safeguards mostly inspect the request before generation or the image after generation, leaving unanswered a mechanistic question: when does a policy- violating edit become internally represented and causally committed? We introduce TRACE-EDIT, a trajectory-level analysis that separates Safety Emergence (SEC), Causal Commitment (CCC), and Visual Evidence (VEC). SEC predicts the realized final unsafe-valid-edit outcome and is evaluated under held-out paraphrase families and same-instruction or different-seed outcomes. CCC is measured with both full- residual and low-rank projected donor interventions plus compatibility controls. Matched counterfactual edits derived from SafetyPairs allow paired trajectories to share a source, edit scope, and random seed. Their temporal ordering defines an Unsafe Commitment Window: a middle region where risk is causally actionable before it becomes visually obvious. We turn this observation into TRAJGUARD, a read-only three-checkpoint monitor with frozen checkpoint, block, threshold, and persistence rules. At the matched 5% benign false-block operating point, the common-protocol endpoint is 19.4% unsafe-valid-edit ASR versus 31.3% for Safety-Guided Flow, with a paired absolute gap of 11.8 pp [4.3,19.4]. A matched three-checkpoint visual monitor remains weaker (28.5% ASR), isolating the value of internal states from earlier temporal inspection alone. The study includes adaptive attackers, independent output/edit judges with human audit, category-held-out analysis, SafeEditBench transfer, and model-native replication on LongCat-Image-Edit and OmniGen2.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.