Actions are Experiments: Identifying Action Semantics for Prediction and Decision
Abstract
The same numeric action can have different physical effects when a robot's controller, coordinate frame, or body changes. We study the resulting ambiguity in action-conditioned prediction: a model must identify what its commands currently mean. We represent this information as an action-semantic state inferred from short response queries, and instantiate it with local affine operators and temporal response kernels. The contribution is a controlled empirical test of this representation, using fixed predictors, held-out interfaces, and paired wrong-state controls. On a joint holdout of UR5e and action programs, after training on Panda and Sawyer, the response state reduces standardized future MSE from 1.7784 to 0.0203; replacing it with a wrong-body state gives 1.8799. A separate probe-then-decide test has zero normalized regret on 72 supported target cases. Official LIBERO-plus and RoboTwin simulations provide further evidence with distinct limits: static and temporal future prediction on RoboTwin both fall short of the prespecified accuracy threshold, while a separate role-goal decision test shows lower endpoint error. We also report the unmet ranking, support, and contact criteria and account for duplicated task dynamics. These results support response-derived action conditioning for tested local role-space futures and decisions. They do not establish benchmark task success, a visual world action model, or a universally sufficient state.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.