Learning Words by Simulating Worlds
Abstract
Result verbs such as open, slice, and fill describe not only what is present in a scene, but how an action transforms one world state into another. We study whether this causal structure provides a privileged learning signal for language models. We train an autoregressive model on interleaved language and visual transition sequences, where each example contains a triple of a pre-condition image, an action description, and a post-condition image. The model is trained with complementary objectives: inferring the action from an observed pre/post transition, predicting the post-condition from the pre-condition and action, and reconstructing the pre-condition from the post-condition and action. Across controlled ablations, visual transition simulation improves grounded word acquisition when it is paired with sufficient language supervision. The gains are concentrated on verbs rather than nouns, and within verbs, on result verbs rather than manner verbs, consistent with the fact that result verbs are defined by observable state changes while manner verbs depend more on motion trajectories. These findings suggest that grounding verb semantics requires more than aligning words with percepts: it requires predictive models of how the world changes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.