acceptodds
Under review as a conference paper at ICLR 2027

CoUPLe: Continuous Flow Unifying Pixels and Language

Abstract

We present CoUPLe, a framework that turns pretrained text-to-image flow models into unified multimodal generators by modeling text responses and images with a shared continuous velocity field. Our key design decouples text conditioning from text generation: a frozen reference encoder and learned connector align conditioning features with the pretrained backbone, while a Text VAE is jointly adapted with the flow model under reconstruction and reference-posterior KL regularization to maintain response decodability. Causal blockwise flow matching enables variable-length responses through autoregressive generation across latent blocks and parallel refinement within each block, with image synthesis treated as a single image-latent block. This formulation supports visual understanding, multimodal reasoning, and instruction-guided editing through interleaved reasoning–image generation. Instantiated on Flux2-9B and Z-Image, CoUPLe supports both MM-DiT and single-stream DiT architectures. Moreover, we find that training CoUPLe with both generation and understanding losses achieves better understanding performance than training with understanding loss alone, demonstrating that generation can help understanding in our settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.