There and Back Again: Bidirectional Diffusion Bridges for Multimodality Translation
Abstract
Multimodality translation (e.g., text-to-image) is a core generative AI task. However, existing approaches (1) have semantically meaningless generative paths, greatly limiting the flexibility of sampling algorithms; (2) are unidirectional, preventing inversion (e.g., image-to-text). We propose BIT: Bidirectional Image-Text Diffusion Bridges. In contrast to previous approaches, BIT starts directly from text and interpolates into images, providing (1) a semantically rich generative path that enables diverse and flexible sampling algorithms; (2) inherent reversibility of the model from image to text, providing a unified, bidirectional generative framework. BIT is derived through rigorous stochastic calculus, yielding SDE forms friendly to simulation and tractable loss functions that scale to high dimensions. Our theoretical and empirical results confirm competitiveness, and often advantage over denoising diffusion and deterministic flow model baselines; not only in vision-language, but also in natural science tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.