acceptodds
Under review as a conference paper at ICLR 2027

Finite-Sample Training, Generalization, and Memorization in Transformer DDPMs

Abstract

Diffusion models can either generate novel samples or reproduce data from their training samples, exhibiting different generalization and memorization behaviors. However, the theoretical mechanisms underlying these behaviors remain poorly understood. Specifically, it is unclear how finite-sample training shapes the learned denoiser and when it leads to memorization. To the best of our knowledge, this paper provides the first finite-sample convergence analysis of Transformer-based diffusion models trained by gradient descent and a theoretical characterization of the resulting sample-specific memorization performance. We consider learning data from a Gaussian mixture distribution under the Denoising Diffusion Probabilistic Model (DDPM) objective with a one-layer single-head Transformer model. We show that gradient descent learns the empirical mode means and within-mode variance, driving the training loss close to the empirical oracle denoising risk, while the generalization gap diminishes with the sample size. We further derive an explicit Gaussian-CDF characterization of a reconstruction-based sample-specific memorization rate, which yields a sample complexity threshold separating sample-specific memorization from generalization. A more fundamental investigation uncovers an empirical-prototype attraction mechanism underlying memorization in diffusion models, i.e., the trained denoiser interpolates between the scaled noisy input and a learned empirical mode prototype, which transitions from an individual training sample to the population mean as the number of samples increases. Numerical experiments validate these theoretical findings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.