acceptodds
Under review as a conference paper at ICLR 2027

Conflux: Self-Refining Unified Diffusion LLM with Cross-Paradigm On-Policy Distillation

Abstract

We propose Conflux, a unified discrete diffusion large language model (DLLM) for visual understanding and generation tasks. Compared with prior unified diffusion models, Conflux introduces two major innovations. First, we propose a unified cross-paradigm on-policy distillation (OPD) framework that successfully transfer the visual understanding ability of autoregressive VLMs and visual generation ability of continuous flow models, to a unified DLLM. Second, Conflux incorporates a novel progressive self-refining diffusion process that gradually corrupts each token across multiple steps during training, as opposed to adopting the conventional masked diffusion formulation that directly corrupts a clean image token to an uninformative mask token. At inference, this design allows Conflux to make soft commit to a token, enable progressive self-refinement for better visual quality. Through extensive experiments, we show that Conflux achieves state-of-the-art performance among unified DLLMs on diverse tasks including visual question-answering, multi-modal reasoning, object grounding, text-to-image generation, and image editing.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.