IntentEdit: Intent-Level Lifelong Editing for VLM Jailbreak Defense
Abstract
Vision Language Models (VLMs) remain vulnerable to multimodal jailbreak attacks. Correcting these vulnerabilities requires safety edits that generalize across attacks sharing a malicious intent and remain effective through subsequent updates. However, learning a separate safety target for each attack does not enforce consistency across different expressions of the same intent. We propose IntentEdit, an intent-level lifelong editing framework that learns one shared textual target and one shared visual target from each pre-grouped intent batch while retaining instance-specific keys. The method jointly incorporates these mappings into a cumulative least-squares objective with constraints on changes along general-knowledge and benign multimodal directions. Recursive updates retain historical mapping constraints through accumulated statistics. We augment two multimodal jailbreak benchmarks with intent groupings to evaluate sequential safety editing. Across the evaluated VLMs, IntentEdit achieves a better safety–utility balance and lower editing runtime than baseline methods. Evaluations on attacks excluded from editing show within-intent transfer, and evaluations of earlier attacks support the safety retention of previous samples in later updates. Ablations support the contribution of shared targets beyond cumulative editing alone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.