acceptodds
Under review as a conference paper at ICLR 2027

Concepts Persist as Directions: Visual-State Auditing of Unlearned Diffusion Models

Abstract

Concept unlearning aims to suppress unsafe, private, or copyrighted concepts in text-to-image diffusion models. Suppression under direct prompting, however, does not establish whether the unlearned model's intermediate visual states can still support concept recovery when the target prompt is excluded from the optimization objective. To examine this question, we introduce Visual-State RECALL (VS-RECALL), an auditing framework that guides latent optimization with a direction defined by a pair of reference states. A positive reference and a matched negative reference define this direction between their U-Net mid-block states under null-text conditioning. VS-RECALL optimizes an adversarial latent by aligning its state displacement from the negative reference with that direction, and requires no access to the original pretrained model during optimization. The optimized latent is then decoded into an image and supplied to the model's image-to-image pathway with the original evaluation prompt. We evaluate VS-RECALL across seven tasks covering unsafe attributes, artistic styles, and object concepts, against ten concept-unlearning methods. VS-RECALL recovers target concepts against all ten unlearning methods and achieves the highest recall success rate, averaged over these methods, on six of the seven tasks, with the largest margins over the strongest respective baselines reaching 10.00% on Golf ball and 8.31% on I2P. These findings show that reference-defined directions in the unlearned model's visual states can guide concept recovery, motivating unlearning evaluations that assess both prompt-level suppression and internal recoverability. Code, prompt sets, and reference pairs are available at https://anonymous.4open.science/r/VS-RECALL.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.