Leveraging Diffusion MLLMs for Image Editing via Dynamic Masked Generation
Abstract
Diffusion Multimodal Large Language Models (dMLLMs) have recently emerged as a powerful paradigm for multimodal generation. Their localized and progressive masked generation property naturally provides a promising foundation for training-free text-based image editing, \ie, explicitly preserving editing-irrelevant regions from the source image, while treating regions to be edited as masked tokens for regeneration guided by the target prompt. Existing dMLLM-based editing method, however, relies on fixed edit masks, static source guidance, and pre-defined generation trajectories, resulting in inflexible edit localization and degraded source consistency. In this paper, we propose DynEdit, a general training-free framework that better exploits the intrinsic generation properties of dMLLMs for flexible and faithful image editing. Specifically, we introduce a parallel dual-branch framework that enables dynamic masked generation with three key designs: 1) Dynamic edit localization to automatically identify editable tokens throughout generation. 2) Adaptive source guidance to adjust source preservation according to specific editing areas and generation steps. 3) Edit-aware token decoding to dynamically refine the decoding schedule for more reliable generation. Extensive experiments on diverse image benchmarks have demonstrated that DynEdit consistently improves editing quality, editing efficiency, and source preservation under both mask-free and external mask-grounded settings, exhibiting strong generalization in various editing scenarios.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.