R2E-Prompt: Reasoning-to-Executable Prompt Generation for Implicit-Instruction Image Editing
Abstract
We present R2E-Prompt, a training-free framework that formulates implicit-instruction image editing as a structured reasoning task in language space. Existing editors translate explicit commands faithfully but struggle with abstract conditions whose visual consequences are not directly named, leaving the entailed changes unrealized in the output. R2E-Prompt addresses this by moving reasoning entirely upstream of the editor: a VLM and an LLM produce, in a single pass, an explicit element-level edit specification consumable by any off-the-shelf instruction-based editor. The framework comprises three stages: Target-State Reasoning first grounds the source image in a source caption and then derives, via chain-of-thought reasoning, an explicit target caption that makes the implicit visual consequences of the instruction explicit; Edit-Aware Scene Graph Transformation parses the source and target captions into text-derived scene graphs and casts source-to-target editing as a structured prediction problem over scene-graph elements; and Executable Prompt Generation linearizes the labeled graph into a natural-language editing command. Without architectural modification or fine-tuning, R2E-Prompt achieves state-of-the-art performance among open-source baselines on RISEBench.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.