M²Change: Captioning and Paired Grounding of Multiple Minimal Changes
Abstract
Grounded image-difference understanding asks models to describe what changes between two images and locate the visual evidence for each change. Most existing benchmarks consider one change at a time, but a pair may contain several changes whose descriptions must be linked to the right objects and regions. We introduce MChange, a benchmark of natural-image pairs with 2–4 simultaneous changes drawn from object, attribute, counting, and relation categories. For each change, models produce a description, entity labels, and boxes in the original and changed images. The test set contains 1,063 pairs and 2,525 annotated changes for evaluating individual predictions and complete-set recovery. We also propose Paired-view Distillation and Set Optimization (PDSO), a three-stage training framework. Supervised fine-tuning first teaches the structured output. In paired-view on-policy distillation, a frozen teacher scores the same response with change-focused and unrelated control crops. This comparison gives greater distillation weight to tokens whose teacher predictions differ between the two views. Finally, set-level reward optimization encourages complete, correctly grounded predictions. Inference uses only the image pair. PDSO reaches 47.2 MCJS, 5.8 points above the strongest evaluated baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.