acceptodds
Under review as a conference paper at ICLR 2027

MUNITE: Unified Multimodal Latent Inference for Any-to-Any Multimodal Generation

Abstract

We introduce MUNITE, a latent-variable framework for flexible any-to-any multimodal generation that treats encoding and latent generation as the same inference problem under different amounts of observed evidence. Given any subset of modalities, MUNITE models the conditional distribution over the latent representation associated with the complete observation. Full observation recovers deterministic encoding, no observation recovers the latent marginal, and intermediate subsets define conditional latent inference, all within a single conditional flow model. A shared latent sample captures variation that must remain consistent across generated targets, while modality-specific generative decoders model the remaining uncertainty independently. To learn these conditional distributions from incomplete training examples, we extend conditional flow matching through self-distillation: predictions conditioned on richer available observations supervise the same model conditioned on smaller subsets at the same intermediate latent state. When the richer-evidence trajectory follows the exact conditional flow, this provides the same expected learning signal as full-target denoising. Across PolyMNIST-D-Q, FFHQ64, and image–text–audio, MUNITE achieves competitive or better generation quality and source–target alignment, with higher joint-generation coherence. In particular, it attains the highest coherence in all one-to-many and unconditional image–text–audio comparisons, showing the effectiveness of unified latent inference across diverse multimodal settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.