acceptodds
Under review as a conference paper at ICLR 2027

Analysis and Mitigation of Memorization in Multi-Modal Diffusion Models

Abstract

Multi-modal diffusion models, such as BLIP-Diffusion and ELITE, achieved impressive performance in generating personalized images conditioned on both a reference image and a textual prompt. Despite their success, it remains unexplored how these models memorize their training data. Existing memorization analyses for text-to-image diffusion models consider the text prompt as the only memorization factor without considering reference image input, making them inapplicable for the multi-modal setting. To address this gap, we propose a novel approach to effectively measure memorization in multi-modal diffusion models. First, our method localizes objects based on caption keywords, then masks each detected region, and measures how faithfully the model reconstructs the masked object. We show that our method provides equivalent results to the gold-standard leave-one-out memorization metric at a substantially lower computational cost. Using our new method, we find that pre-training samples for the multi-modal encoder exhibit consistently higher memorization levels than samples used to fine-tune the diffusion backbone. This indicates that the multi-modal encoder plays a more important role in driving memorization than the diffusion backbone itself. We further demonstrate that highly memorized training images are under high privacy leakage risks. This is driven by the joint memorization of vision–text modality associations, where the vision modality matters more. Finally, we show that applying memorization mitigation at the multi-modal encoder level achieves a significantly better privacy-utility tradeoff than diffusion backbone level mitigation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.