acceptodds
Under review as a conference paper at ICLR 2027

Tangible and Editable Causal Action-to-Object Pathways in Robot World Models

Abstract

Robot world models can predict action outcomes yet still learn questionable action-to-object dependencies, making their predictions difficult to trust and their failures difficult to diagnose or correct. Existing black-box transitions obscure how the current action causally contributes to an object's predicted change within the learned model, while post-hoc attribution does not expose a separate causal pathway that can be directly revised. We address this limitation by making the model-internal causal contribution of the current action an explicit, auditable, and editable component of the world model. This enables potentially unreasonable learned action effects to be examined against independent simulator interventions and locally modified without retraining the entire dynamics network. To achieve this, we introduce a state-conditioned causal action-to-object pathway that participates directly in the learned transition and separates the current action's contribution from the remaining object dynamics. The pathway is learned jointly with the world model without direct action-effect or counterfactual supervision. We further develop a minimum-change causal mechanism edit that assigns the same executed action a prescribed alternative contribution to the predicted object transition while keeping the current state and action fixed. Experiments on Meta-World manipulation tasks show that the exposed pathway faithfully reflects the model's forward computation, aligns with independently measured action influence in the simulator, and supports precise targeted editing. Separate closed-loop interventions further demonstrate that modifying the learned pathway changes robot behavior and can yield exploratory success gains in selected tasks without retraining. These results show that world models can be designed not only to predict the consequences of actions, but also to expose, audit, and locally revise the internal causal mechanisms through which those consequences are represented, providing a foundation for assessing the reliability of opaque world models and exploring targeted interventions in failing tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.