Immune Escape in Model Immunization
Abstract
Efficient fine-tuning techniques have substantially reduced the cost of adapting open-weight models, enabling malicious actors to readily fine-tune them toward harmful capabilities. In response, model immunization modifies a pretrained checkpoint to resist subsequent adaptation toward designated harmful capabilities while preserving its utility and adaptability for benign tasks. However, we observe that existing immunization methods substantially overfit to the specific configurations used during immunization, particularly the images, textual prompts, and adaptation strategy the surrogate attacker is assumed to use. This fragility is obscured by existing evaluations, which largely reuse the same datasets and attack configurations employed during immunization. To rigorously evaluate model immunization, we develop ESCAPE, an attack framework that generates diverse variants of harmful adaptation designed to circumvent immunization. Specifically, ESCAPE misaligns the attacker from the configuration simulated during immunization along three axes: harmful data, model initialization, and adaptation strategy. We evaluate ESCAPE against three existing model immunization methods across diverse tasks and modalities, spanning text-to-image generation, image classification, and language models. Our results show that ESCAPE substantially weakens immunization across all three modalities and, in many cases, largely erases the protection that immunization provides over the un-immunized model.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.