PGAMIE: A PIPELINE FOR GENERATING ACCURATE MULTI-INSTANCE IMAGE EDITS
Abstract
Text-guided image editing has advanced rapidly, enabling increasingly complex image modifications through natural-language instructions. Despite this progress, models often struggle to add or remove an exact number of object instances, particularly when instructions involve multiple object categories or edit operations. We present the first systematic evaluation of multi-instance editing and introduce , an automated pipeline for constructing verified multi-instance, multi-category image-editing data. The pipeline generates count-verified scenes and constructs verified addition and removal edits through instance segmentation, inpainting, and agreement-based count verification, supporting single-category, multi-category, and compound edits in scenes containing up to 44 object instances. Using this pipeline, we create MOMI, a 20k-record training dataset and a multi-category benchmark, together with two complementary benchmarks for out-of-distribution single-category and real-image evaluation. Our evaluation reveals that multi-instance editing remains challenging for current image editors, with performance degrading sharply as requested object counts and edit complexity increase. Fine-tuning Flux Kontext on MOMI achieves the highest edit accuracy on the MOMI benchmark at 16.3%, compared with 2.0% for the base model and 14.3% for Nanobanana Pro.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.