acceptodds
Under review as a conference paper at ICLR 2027

Orthogonality of Generalization and Memorization Trajectories in Diffusion Models

Abstract

Training over-parameterized diffusion models with limited data typically exhibits a two-phase phenomenon in which the model generalizes early in training before progressively memorizing the training data. Although this qualitative behavior has been observed and studied in many recent works, a precise quantitative characterization of the geometry and dynamics underlying this transition is still lacking. In this paper, we characterize this transition through the geometry of the denoiser outputs. We show that the training trajectory can be decomposed into two directions associated with the population and empirical distributions, which are orthogonal in expectation. Across synthetic and real-world datasets, this decomposition reveals a quantitative phase-transition phenomenon with the denoising trajectory: the trajectory first moves predominantly toward the population Bayes-optimal denoiser, then it makes a sharp, nearly orthogonal turn toward the empirical Bayes-optimal denoiser as memorization emerges. We further study this behavior through the spectrum of the denoiser output-space training kernel, observing that the population direction is learned through substantially faster kernel modes than the empirical direction. Motivated by this separation, we propose a Direction-Aware Orthogonal (DAO) regularization, which selectively penalizes denoising motion along an estimated empirical direction. Experiments show that DAO is highly effective for mitigating memorization while further improving generation quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.