acceptodds
Under review as a conference paper at ICLR 2027

Actions Without Consequences: Binding World Action Model Policies to What They Decide

Abstract

World action models (WAMs) build on pretrained video models to jointly predict future observations and robot actions. Demonstration data, however, provides only one action for each observed state and therefore does not directly constrain how the predicted future should change under alternative actions. We study this gap with CF-BENCH, a counterfactual benchmark for measuring action sensitivity in WAM policies. \anchor is introduced to compare the robot pose represented in the predicted future with the pose implied by the predicted action through forward kinematics and camera calibration. This connection also enables counterfactual training without ground-truth future observations: we perturb the action, enforce the corresponding kinematic change in the robot region, and constrain regions outside the arm's swept volume to remain unchanged. With counterfactual training, ANCHOR improves the command-normalised keypoint response to 0.79 for in-distribution perturbations and 0.71 for held-out perturbations. Under matched evaluation settings, it also improves task success over base-FT by 2.3–6.6 percentage across RoboCasa, RoboTwin 2.0, and LIBERO-Plus.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.