acceptodds
Under review as a conference paper at ICLR 2027

Let Relations Shape the World: Learning Instruction-Conditioned World Models

Abstract

Relational instructions can express the same task across different participants, yet demonstrations provide only scene-specific visual outcomes. We introduce RWM, a relation-shaped world model that learns language-grounded object interactions to predict desired futures. Its relational transformer, RELATRON, grounds language in ordered object pairs through a read–write operation, with separate branches determining which interactions contribute and what information they convey. Both branches are trained end-to-end with the future predictor using demonstrated visual changes, without explicit object-pair relation annotations. Shared pairwise operators provide an inductive bias for reuse across participants, while their spatial context augments rather than replaces dense visual features. The predictor generates short-term and terminal visual goals supervised by future observations rather than symbolic goal-state targets. These goals guide action proposals, whose predicted consequences are compared with the short-term goal in a shared visual representation space. Controlled evaluations on our TELETRAAN I benchmark show higher future-goal agreement under simultaneous target–reference color shifts, improved instructed-object grasping under joint target–distractor changes, and higher mean pickup success on colors excluded from task training. Ablations show contributions from both interaction selection and language-conditioned messages. On LIBERO, consequence-based selection improves mean success over single-proposal execution. These results support learning relational conditioning from demonstrations to apply familiar task requirements under changes in participant appearance and scene context.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.