LIFT-WAM: From Counterfactual Adaptation to Grounded Goal Execution
Abstract
In language-guided manipulation, the same scene can support multiple feasible goals, with the current instruction specifying the intended outcome. World-action models (WAMs) may continue executing scene-associated behaviors after an instruction changes, and correct target interaction can still fall short of establishing the requested final relation. We characterize this gap at three levels: alternative-goal responsiveness, goal grounding, and complete goal realization. Our counterfactual post-training framework, LIFT-WAM, learns alternative-goal behavior from native and counterfactual supervision. A Grounded Intent Extractor and an Intent-to-Action Adapter convert scene-grounded entity roles and target relations into explicit action conditions, while replay-verified corrective continuations supervise execution through goal completion. Under same-scene counterfactual evaluation with source initial states held fixed, LIFT-WAM achieves success on LIBERO-10/LIBERO-Object and on RoboTwin Clean/Random, compared with and , respectively, for counterfactual adaptation alone. Controlled Video Expert interventions further show that, after counterfactual adaptation, language-conditioned representations of current observations can redirect closed-loop control.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.