Paired Flow Matching: Representation Learning without a Reconstruction Bottleneck
Abstract
Variational autoencoders (VAEs) are a standard tool for representation learning, but their samples often have lower fidelity than diffusion and flow models that generate directly in observation space. Latent diffusion and latent flow matching reduce computation by generating in a VAE latent space, yet they too lose sample quality. The shared limitation is the encode-and-decode design: a compact representation should discard presentation details, but the decoder needs those details to reconstruct the observation. The representation must therefore either retain unwanted detail or incur reconstruction error. To resolve this conflict, we propose Paired Flow Matching (PairedFM), which uses representations to guide transport between full observations. From ordered pairs drawn from a single data distribution, PairedFM jointly learns a representation and a family of flows in observation space, indexed by the pair's representations. At generation, a source observation supplies the details and a target representation determines where to move. PairedFM also comes with theoretical guarantees: We characterize when its representation is identifiably and the resulting transport achieves exact recovery of the truth. Across empirical studies, we find that PairedFM learns informative representations while maintaining sample quality where VAE and latent flow matching baselines degrade.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.