acceptodds
Under review as a conference paper at ICLR 2027

Learning Minimal Sufficient Visual Intervention for Instruction-Based Image Editing

Abstract

Recent advances in instruction-based image editing have enabled high-quality edits from source images and language instructions. However, the visual intervention needed to translate user intent into executable editing conditions typically remains implicit, leaving the intended change boundary underspecified and potentially leading to incomplete edits or unintended modifications to unrelated content. In this work, we formulate instruction-based image editing as learning a minimal sufficient visual intervention (MSVI) that fulfills the instruction while minimizing changes to instruction-invariant visual content. Specifically, we first use a multimodal large language model as an edit planner to predict a structured edit specification encoding the intervention regime, visual transition, semantic actions, and spatial layout. Based on the predicted specification, a deterministic compiler then invokes SAMĀ 3 only when the localized grounding is required. The resulting outputs form the intervention condition supplied to a flow-matching editor. To supervise the planner and the editor, we further construct MSVI-13K with annotated structured edit specifications. The annotations provide prediction targets for the planner and determine the intervention conditions used to train the editor with conditional flow matching and a source-copy suppression loss. Extensive experimental results on MSVI-13K and two external benchmarks show that our method matches or outperforms state-of-the-art approaches, supporting explicit visual intervention learning for balancing instruction fulfillment and content preservation. Code is available at https://anonymous.4open.science/r/MSVI/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.