CONFLUX: Role Separation for Unified Multimodal Modeling
Abstract
Unified multimodal models with a shared backbone face costly backbone iterations for image generation and conflicting visibility requirements for prediction and conditioning. We introduce CONFLUX, which separates semantic conditioning from rendering and observed content from prediction queries. The backbone com- putes image conditions once per target image and conditioning branch; a diffusion transformer reuses them for iterative rendering. We adopt content–query factor- ization across modalities with a common representation interface: queries encode leakage-free prediction conditions, while content states preserve complete, clean image context for subsequent predictions. With a fully trainable 0.6B backbone and matched data and update budgets, CONFLUX improves MMLU by 7.56 points and ImageNet classification by 3.53 points over the strongest respective baselines in the main comparison. Image generation achieves 4.0× the throughput of the fastest baseline under common settings, with comparable FID (3.67 versus 3.09) at separately selected generation configurations. Scaling only the flow head improves FID to 3.19 while largely preserving language and visual understanding performance and retaining a throughput advantage.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.