VEGA: Sparse Mark Experts for Composable Visual-Instruction Editing
Abstract
Precise image editing has long sought spatial instructions that express not only what to change, but also where to change. Modern multimodal editors are strong at understanding semantic intent from language, yet precise spatial grounding remains difficult to express through text alone. Visual instructions, such as boxes, arrows, and doodles, provide this missing grounding directly in image coordinates, while existing editors struggle to execute such instructions reliably. We study this failure and find that visual marks are not simply inaccessible to the editor. A frozen editor can localize and manipulate these marks when they are treated as image content, yet often fails when the same marks serve as editing instructions. This gap suggests that existing editors have not learned to reliably use visual marks as editing controls. We propose VEGA, a lightweight adapter that turns visual instructions into executable editing controls. For each type of visual instruction, VEGA trains a specialized expert using flow matching with auxiliary attention supervision, encouraging the expert to route the corresponding spatial evidence into localized editing responses. Compound training further encourages each expert to remain selective in the presence of competing annotations. We then introduce Attention Response Routing, which uses these learned responses as token-wise responsibilities to compose independently trained experts along a shared generation trajectory, without parameter merging or joint expert training. Experiments on internal and external benchmarks show that VEGA is effective across multimodal editing models, substantially improving spatially grounded editing and enabling efficient composition of multiple visual instructions while preserving the original generation quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.