acceptodds
Under review as a conference paper at ICLR 2027

Weighting Schedules Govern What and When Score-Based Generative Models Learn from Multimodal Data

Abstract

Score-based generative models generate new samples by integrating a time-dependent drift that carries Gaussian noise onto the target distribution. In practice this drift is modeled by a neural network, trained on a loss integrated over time with a weighting schedule . Along the backward dynamics, and for multi-modal distributions, trajectories commit to modes of the target within a narrow time window, the speciation time. In this work, focusing on high-dimensional data, we decompose the integrated loss into its single-time contributions and analyze each at fixed signal-to-noise ratio : we show that sets the rate at which each feature of a multimodal target—the mode directions and their relative weights—is acquired during training. Crucially, at high all mode directions are acquired together, on a single timescale insensitive to their amplitudes, while the relative weights are not learned at all. Only near the speciation time, where becomes of order one, do all features become learnable, each on its own timescale: the weights are acquired jointly with the directions, and the directions at rates set by their relative amplitudes. For models trained on time-integrated objectives, the learning dynamics is then governed by how much of the weighting effectively sits near the speciation time, which provides insights on design choices. These results follow from an exact high-dimensional analysis of the training dynamics of unbalanced and hierarchical Gaussian mixtures. Numerical experiments on image and human genome haplotype generation recover the predicted hierarchy of learning timescales in more complex settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.