acceptodds
Under review as a conference paper at ICLR 2027

Diagnosing and Mitigating Memorisation in Generative Models with Metric Space Magnitude

Abstract

Memorisation in deep generative models is often reduced to the question of whether a similarity threshold is crossed. Current successful approaches typically analyse sample-wise comparisons between generated outputs and individual training examples. In this work, we propose a framework based on metric-space magnitude, an invariant that measures the effective number of distinct samples across multiple scales, providing a notion of diversity. For example, in an image embedding space, similar images contribute less effective diversity than distinct ones. By comparing the effective diversity of the training and generated sets and their combination, we shift from pointwise copy detection to a set-level description of how strongly and in what form the generated distribution is absorbed into the training set. We apply our framework to image, 3D, and neural-network parameter generation, where it provides a richer quantitative account of how memorisation is expressed in these settings, for example, whether it is broadly distributed over the training set, concentrated in a small part of it, or accompanied by self-collapse. This broader view remains consistent with the main findings of previous analyses. Finally, we use our approach to construct post-training interventions for image and face generation, substantially reducing memorisation while avoiding undesirable geometric side effects.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.