acceptodds
Under review as a conference paper at ICLR 2027

Seeing What Changed: Embodied Change Grounding with Selective Cross-View Perception Optimization

Abstract

Embodied agents often revisit the same environments, where objects may appear, disappear, move, or change state between visits. To study how an agent can tell what has changed and where, we formalize Embodied Change Grounding (ECG): given a reference image, a query image captured after the scene has changed, and a language instruction, a model must identify the referred change and segment the changed object, which may lie in either image. As the two egocentric views are unaligned, this requires cross-view correspondence before localization. We introduce ECG-Bench, a benchmark of 10K samples from 1,859 simulated, industrial, and real revisited scenes with six change types, and ECG-50K, a training set with Change-CoT traces. Our evaluation shows that current MLLMs, reasoning segmentation frameworks, and RL post-training methods all fall short on ECG, and that models often produce fluent reasoning that is not grounded in the cross-view evidence. We therefore propose ECG-Grounder, whose CoT-Guided Cross-View Decoder selects the view containing the target, aligns the two views, and decodes the mask under the guidance of the reasoning chain. We further introduce Selective Cross-View Perception Optimization (SCPO) to train the model, rewarding reasoning chains whose outputs are more sensitive to removing the change-critical region than to a matched corruption elsewhere, as measured at high-entropy tokens within GRPO. ECG-Grounder with SCPO achieves the best results on ECG-Bench and also generalizes to other multi-image reasoning benchmarks such as MIG-Bench.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.