MoG-Diff: Mixture-of-Gaussians Latent Structuring for Controllable Diffusion Bridge Generation
Abstract
Diffusion-bridge autoencoding connects representation learning, reconstruction, and generation, yet its latent space need not be semantically organized or provide explicit coordinates for traversing generative modes. We introduce MoG-Diff, a unified framework that structures the latent space with a mixture of Gaussians and represents each sample through a component and a component-relative residual. Variational mixture inference learns the component organization, while a task-adaptive interface connects the structured latent to the endpoint-conditioned bridge. A pointwise Doob h-transform analysis establishes compatibility with the bridge construction. The resulting geometry supports component-conditional generation, residual-preserving and responsibility-controlled pairwise transport, and multi-component responsibility simplex control. On MNIST and Fashion-MNIST, MoG-Diff achieves the best clustering metrics among the evaluated methods, complete component-majority class coverage, and the lowest Conditional Feature FID. On CelebA, it performs on par with or slightly better than DBAE in reconstruction and provides qualitative natural-image extensions of component generation and transport. Responsibility control reduces pairwise MAE from 0.2065 to 0.0356 and three-component simplex MAE from 0.0644 to 0.0227. These results show that Gaussian-mixture structure can serve not only as a representation prior but also as an operational geometry for controllable diffusion-bridge generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.