acceptodds
Under review as a conference paper at ICLR 2027

M²Change: Captioning and Paired Grounding of Multiple Minimal Changes

Abstract

Grounded image-difference understanding asks models to describe what changes between two images and locate the visual evidence for each change. Most existing benchmarks consider one change at a time, but a pair may contain several changes whose descriptions must be linked to the right objects and regions. We introduce MChange, a benchmark of natural-image pairs with 2–4 simultaneous changes drawn from object, attribute, counting, and relation categories. For each change, models produce a description, entity labels, and boxes in the original and changed images. The test set contains 1,063 pairs and 2,525 annotated changes for evaluating individual predictions and complete-set recovery. We also propose Paired-view Distillation and Set Optimization (PDSO), a three-stage training framework. Supervised fine-tuning first teaches the structured output. In paired-view on-policy distillation, a frozen teacher scores the same response with change-focused and unrelated control crops. This comparison gives greater distillation weight to tokens whose teacher predictions differ between the two views. Finally, set-level reward optimization encourages complete, correctly grounded predictions. Inference uses only the image pair. PDSO reaches 47.2 MCJS, 5.8 points above the strongest evaluated baseline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.