acceptodds
Under review as a conference paper at ICLR 2027

Your Unlearning Gives You Away: Identifying Erased Concepts in Diffusion Models

Abstract

Existing attacks on unlearned diffusion models assume that the erased concepts are known in advance and focus on recovering them. In practice, however, model providers may not disclose which concepts have been removed, and even with access to the original base model, an adversary may still lack a clear target to attack. In this paper, we aim to answer the following critical but overlooked questions: ***which concepts have been erased from the model, and how many have been erased in total?*** To this end, we present TRACER, a framework that rapidly and accurately identifies erased concepts and estimates their number. TRACER efficiently identifies erased concepts without generating and classifying images. Using lightweight spectral analysis of weight footprints, it enables efficient search over large candidate vocabularies. To distinguish multiple erased concepts, we introduce a footprint coverage objective that guides sequential discovery. TRACER estimates the number of erased concepts by detecting a sharp decline in candidate confidence as the selected concepts account for the erasure footprint, without requiring labeled examples for calibration. The framework requires only lightweight linear algebra and limited forward probes, with no prior knowledge of the unlearning algorithm. Experiments across text-to-image and text-to-video backbones and diverse unlearning methods demonstrate that TRACER identifies erased concepts and estimates their number *in seconds*, achieving **150–137,000×** and **133–20,000×** speedups over MIA and brute-force search on image and video models, respectively, with substantially higher identification accuracy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.