acceptodds
Under review as a conference paper at ICLR 2027

WorldEffect: Learning Action-Conditioned Object Dynamics from Human Plans

Abstract

Existing world models primarily predict future observations, but embodied systems must additionally infer how a planned human action changes the physical state of objects in the environment. This problem becomes difficult when nearly identical motions lead to different object responses, as when a hand touches an object or narrowly misses it. Sparse contact can be obscured by shared scene motion, the coordinates that best describe an interaction change with the grasp, and even a specified plan can admit multiple outcomes. We introduce WorldEffect, an action-conditioned object-dynamics framework that predicts object responses from observed history and a supplied human plan. WorldEffect learns which aspects of a planned action induce object changes by comparing responses to plans with the same history. It combines coordinate views that separately assess translation and rotation, then assigns plan-conditioned probabilities to a fixed set of plausible outcomes without seeing the target response. On the BEHAVE dataset, WorldEffect reduces mean object translation error by 16.01% over fixed contact transport, which moves the object according to the planned human motion. For the same 16 candidate trajectories, its input-conditioned ranking reduces probability-weighted joint trajectory error by 12.73% relative to uniform weighting. Together, these results show that explicitly modeling action-induced object responses improves both deterministic prediction and uncertainty-aware future ordering.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.