acceptodds
Under review as a conference paper at ICLR 2027

Memorization Detection in Diffusion Models via Text Embedding Interpolation

Abstract

Text-to-image diffusion models generate high-fidelity images from a text prompt but exhibit memorization on certain prompts where the model reproduces a training image. A widely adopted approach to memorization detection quantifies the difference between the conditional and unconditional scores. We express the magnitude of this difference as the norm of the integral of the conditional score's derivative along the linear interpolation from the unconditional embedding to the prompt embedding. For memorized prompts with an anomalously large score difference, the rate of change of the conditional score remains relatively stable over most of the interpolation path and rises sharply within a narrow interval near the prompt embedding. The generated image follows the same trajectory remaining aligned with the prompt semantics across the smooth segment and switching to the memorized training image across the narrow interval. We measure this sharp rise through the condition Jacobian evaluated at the prompt embedding which yields a detection performance that achieves state-of-the-art accuracy. Detection performance remains stable over a range of reduced latent resolutions, allowing the resolution to be selected to reduce memory consumption with minimal performance loss.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.