acceptodds
Under review as a conference paper at ICLR 2027

FLUX Learns To Write

Abstract

Multimodal Diffusion Transformers (MM-DiTs) jointly attend to image and text tokens, introducing image-to-text feedback and implicit visually grounded semantics, making them promising for unified understanding and generation. However, existing approaches largely expose this capability only through pixel-level grounding and editing, while their reliance on full-sequence text denoising limits coherent long-range reasoning over jointly grounded visual and textual information. Moreover, weak cross-modal coupling limits visual feedback to textual reasoning, leading to semantic misalignment. We propose UniFlux, which adapts a pretrained MM-DiT for multimodal understanding using block-wise discrete flow matching, for long and coherent multimodal reasoning while preserving its generation prior. We further introduce X-GRPO, a cross-modal reinforcement-learning objective that couples separately normalized text- and image-side advantages through positive-gated multiplicative credit, assigning positive rollout credit only when both reasoning quality and image-regeneration fidelity exceed their respective group baselines. UniFlux achieves state-of-the-art performance across eleven understanding and generation benchmarks

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.