acceptodds
Under review as a conference paper at ICLR 2027

ControlEdit: Image Editing with Unpaired Post-Training

Abstract

Recent instruction-based editing models achieve impressive editing performance by fine-tuning pretrained text-to-image models with large-scale pair (source image, target image) editing data. However, constructing such pair editing data is expensive, often requiring a complex data pipeline and human verification, which makes the process of finetuning text-to-image model to image editing model costly and difficult. In this work, we propose \method, a post-training framework that converts a pretrained text-to-image model into an instruction-based image editing model without requiring target image during training. Our framework consists of two stages. The first stage converts text-to-image diffusion model to take instruction and source image as editing condition, while the second stage applies reinforcement learning post-training to further optimize editing quality with respect to prompt alignment and background preservation. Experiments across multiple editing tasks demonstrate that \method achieves competitive performance with supervised editing methods while eliminating the need for paired edited images.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.