Efficient Extraction Attacks for Unconditional Diffusion Models
Abstract
Diffusion models (DMs) have become one of the most powerful image generation models. However, the data privacy risks associated with them have not been fully revealed. The data leakage problem for unconditional DMs trained on images of human faces is especially concerning. In this paper, we ask: "Is it possible to extract training data efficiently from the DMs pretrained on human face datasets?" To answer the question, we first formulate the task of efficient extraction attacks in the context of human face images. We first show efficiency of extraction attacks depends on both the speed of generation and per-sample hit probability. By testing mainstream sampling acceleration methods including SDE and ODE solvers, we validate that faster samplers are indeed more efficient in extracting training data, but ODE solvers are still limited by the low per-sample hit probability. Although the SDE solvers are computationally more intensive per image, they are five times more likely to extract a training image with each generation, making them much more efficient than the ODE counterparts. This suggests that hit probability is a more critical determinant of efficiency than generation speed alone. To fundamentally boost sampling efficiency by increasing the hitting rate, we propose a Generation-Filtering-Resampling framework that adapts the distribution of initial noise to concentrate on the high-probability regions in the latent space. Experiments on both a pixel-space DM and a latent-space DM trained on CelebA-HQ and FFHQ validate the efficiency improvement of our method, which is at least twice as efficient as the existing methods in most cases.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.