acceptodds
Under review as a conference paper at ICLR 2027

GIFT: Gradient-aware Immunization of diffusion models against malicious Fine-Tuning with safe concepts retention

Abstract

Text-to-image diffusion models can learn semantic classes, like artistic styles or object categories, which model owners may need to remove to prevent harmful or unauthorized outputs. However, existing concept-erasure methods are brittle; adversaries can easily recover the erased concepts using simple fine-tuning attacks. This vulnerability occurs because erasure methods only increase the loss gap for the target concept but leave the underlying representational structure intact, which makes the concept easily "reachable" during fine-tuning. To address this, we introduce GIFT, an immunization technique designed to attack both the loss landscape and the hidden representation of a target concept, leading to a tamper-resistant model. GIFT pairs gradient ascent on the denoising loss with representation noising, a technique that pulls intermediate attention activations toward Gaussian noise to degrade the reachability of the target concept. To maintain general model utility on safe concepts, GIFT anchors internal representations to a frozen base model and utilizes gradient surgery to resolve optimization conflicts between the erasure and preservation updates. Extensive evaluations across art styles and object categories demonstrate that GIFT achieves superior initial erasure and remains significantly more robust against fine-tuning attacks compared to state-of-the-art baselines, all while maintaining fidelity on safe-concept generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.