Change2Act: Object-Centric Visual Reasoning for Image-Goal Manipulation
Abstract
Robot manipulation is a fundamental capability for embodied intelligence in real-world environments. Recent generalist robot policies typically specify task goals through natural language and generate continuous actions from visual observations. However, specifying tasks in complex scenes requires precise instructions about objects and their spatial relationships, and writing such instructions demands substantial human effort. To this end, we introduce Image-Goal Manipulation (IGM), where a goal image directly specifies the desired environment state and the policy generates continuous manipulation actions from the current visual observation to realize that state. IGM poses two key challenges. First, the policy must identify task-relevant objects and their desired state transitions despite substantial visual differences between the current observation and goal image. Second, these transitions must be converted into effective action conditions that guide object interaction and continuous control. To address these challenges, we propose Change2Act, a framework that connects explicit visual change reasoning with lightweight policy adaptation. Specifically, we design an Object-Centric Change Reasoning module that establishes semantic correspondences between objects in the initial and goal images and encodes their spatial differences into a structured change descriptor. We further design a Change-Guided Action Conditioning module that maps this descriptor into compact goal tokens and combines them with the original goal image to guide a pretrained robot policy. A goal-role embedding distinguishes the desired state from current observations. Experiments demonstrate that Change2Act achieves more accurate offline action predictions than the compared baselines while requiring fewer trainable parameters.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.