PatchRecover: Image-Centric Diffusion Memorization Detection via Local Counterfactual Patch Recovery
Abstract
Diffusion models can memorize parts of training images, posing privacy and copyright risks. Many existing detectors rely on global, prompt-conditioned signals, which can obscure local memorization and limit detection when only generated images are available. We propose *PatchRecover*, a framework that extracts local evidence through counterfactual patch recovery and can operate without the original prompt, seed, or denoising trajectory. The core idea is that memorized content should recover more readily after a local edit than non-memorized content. We test this by replacing a patch with a counterfactual alternative, regenerating the image, and measuring recovery toward the original content. We aggregate the strongest patch-level recovery margins to obtain an image-level detection score. To explain why memorized patches recover more strongly, we introduce local recovery regions and derive a reconstruction bound under a local contraction assumption. Trajectory measurements show that memorized patches enter these regions more often and recover more strongly than non-memorized patches. On Stable Diffusion v1.4 and v2-base benchmarks that include non-memorized outputs from memorization-associated prompts, PatchRecover achieves mean image-level AUCs of 0.9980 and 0.9968, respectively. When both memorized and non-memorized images come from memorization-associated prompts, PatchRecover improves mean AUC from the strongest baseline's 0.9374 to 0.9782.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.