ActionTrace: Learning Action Consequences in Robot World Models
Abstract
Robot world models offer a way to evaluate policies through imagined rollouts, provided that their predictions reflect the consequences of the supplied actions. However, from the same initial scene, a model may generate similar, visually plausible futures for actions with different physical outcomes, obscuring differences between policies. To preserve these distinctions, we introduce ActionTrace, a robot world model that learns action consequences through spatial and temporal conditioning. It projects joint trajectories into the camera view to guide predicted robot motion and carries an action-updated consequence state across video segments, allowing earlier actions to inform later predictions. We reinforce this action dependence during training by favoring executed controls over alternatives from the same simulator reset for each observed future. Under paired action interventions, ActionTrace predicts object outcomes with higher overall accuracy across recursive rollouts. Across four newly trained policies on 50 manipulation tasks, ActionTrace recovers their ordering with a Pearson correlation of 0.9805 between predicted and simulator success rates. On the WorldArena benchmark, it also recovers the ordering of five released policies, reaching a Pearson correlation of 0.9920. The model achieves these policy-evaluation gains while retaining comparable long-horizon video fidelity. Adapting its components to Ctrl-World extends these policy-evaluation improvements to real-world outcomes across seven tasks and three policies.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.