acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Scribble-Guided Image Editing: A Progressive Supervision Curriculum for Coherent Multi-Task Editing

Abstract

Scribble-guided image editing combines text with coarse strokes to specify both what to edit and where. However, current image editors often fail to follow scribble instructions, especially for multiple requests. We diagnose two gaps in the required training supervision. First, precise supervision and visually coherent supervision play complementary roles: the former supports reliable local editing, while the latter improves visual coherence with the source image. Second, exposure to multiple scribbles alone is insufficient for multi-task editing; complete multi-task supervision is needed for reliable binding between scribbles, instructions, and edits. Based on these findings, we propose a Progressive Supervision Curriculum. Layered Synthetic Data provide precise supervision for single-task edits. Multi-Task Mosaicking extends these samples into complete multi-task tuples. Visually Coherent Synthetic Data provide supervision for visual coherence. Stage I learns single-task editing and multi-task binding from Layered Synthetic Data and mosaicked tuples. Stage II refines visual coherence with Visually Coherent Synthetic Data. We render scribbles directly into the input image and train only a LoRA adapter. On VIBE, our method achieves the best overall single-task result and reaches performance comparable to closed-source models on multi-task editing. We will release our dataset and model.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.