acceptodds
Under review as a conference paper at ICLR 2027

EditThinker: Enhancing Image Editors via Iterative Reasoning with a Plug-and-Play VLM

Abstract

Image editors can produce realistic images while still failing to satisfy the requested changes. Improving their instruction following requires identifying these failures and translating visual feedback into instructions that the editor can execute. We introduce EditThinker, an 8B vision-language model (VLM) that enhances existing image editors through iterative reasoning while keeping their weights fixed. EditThinker jointly critiques semantic consistency and visual quality, reasons about editing failures, and refines the instruction in a structured response. It interacts with the editor through a Critique-Refine-Repeat loop, providing a plug-and-play interface for feedback-driven editing. Learning effective refinements requires more than imitating plausible expert suggestions: the instructions must work with the capabilities of the underlying editor. We therefore combine supervised fine-tuning on expert demonstrations with reinforcement learning from executed editing outcomes, using format, critic, and edit rewards. To support this training, we construct THINKEDIT-167k, a dataset of 167k step-wise samples with unified supervision for critique, reasoning, and instruction refinement. Experiments on four benchmarks show improved overall instruction-following scores across the tested editors. Ablations examine the joint model design, data allocation, and training across VLM backbones. Our round-budget study further shows a favorable quality-editor-call trade-off relative to best-of-N selection. We will release our data construction framework, datasets, and models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.