FlowSeed: Decoupled Source Learning for Cross-Modal Generation
Abstract
Flow-based generative models typically start from condition-independent Gaussian sources, leaving source initialization largely unexplored. We show that source initialization substantially affects generation quality, while jointly learning condition-dependent sources and generators can undesirably entangle source learning with generative modeling. We therefore propose FlowSeed, a decoupled framework that learns condition-dependent sources independently of the image generator. FlowSeed first encodes paired cross-modal correspondence in an aligned representation space, from which independently trained flows recover compatible source structures. For text-to-image generation, a lightweight text flow provides condition-dependent initialization for the aligned semantic component, while a token-wise FlowEdit-style strategy bridges it to the image generation flow. By decoupling source learning from generation, FlowSeed preserves the efficient and diverse generation paradigm of standard Flow Matching while enabling more effective condition-aware initialization. Experiments across multiple representation spaces consistently demonstrate improved generation performance over standard Flow Matching and learned-source baselines. Moreover, FlowSeed supports plug-and-play adaptation of pretrained generators to new conditioning modalities without retraining the generator.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.