Generative Auto Encoders via Picard Iterated Wasserstein Gradient Flows
Abstract
Variational Auto-Encoders are trained with the goal of reconstructing the inputs from the latents accurately, and they also attempt to match the latent distribution to a desired prior (typically standard Gaussian). Due to the tradeoff between reconstruction and prior regularization in ELBO-based training, the latent distribution is far from the desired prior, which results in poor quality of VAEs when used as generative models. Motivated by the problem of shaping a latent distribution to the Gaussian, we introduce the smoothed KL divergence, a geodesically non-convex functional on the space of probability distributions and consider its Wasserstein Gradient Flow (WGF). This WGF requires only the score function of the smoothed distribution and is easily learned via the popular denoising score matching. We establish theoretically that it converges to the target distribution (Gaussian) via a Lyapunov decay argument, bypassing Polyak-Lojaciewicz type conditions usually used in this context. We apply the smoothed KL WGF to derive a novel generative modeling algorithm, based on autoencoders. Our algorithm introduces a time-parameterized neural network in the latent space to represent trajectories, and during training we iteratively update the entire family of trajectories using Picard's iteration, with the goal of eventually aligning the trajectories to the WGF. Simultaneously, we adapt the decoder to reconstruct from the end-points of the trajectories using autoencoder reconstruction losses. We establish theoretical guarantees for convergence of the Picard iterates to the chosen WGF, as well as convergence of the WGF to Gaussian, under mild assumptions on the initial latent distribution. We empirically establish the efficacy of our method on the unconditional datasets CelebA-HQ, LSUN Church and LSUN Bedroom.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.