acceptodds
Under review as a conference paper at ICLR 2027

Learning Words by Simulating Worlds

Abstract

Result verbs such as open, slice, and fill describe not only what is present in a scene, but how an action transforms one world state into another. We study whether this causal structure provides a privileged learning signal for language models. We train an autoregressive model on interleaved language and visual transition sequences, where each example contains a triple of a pre-condition image, an action description, and a post-condition image. The model is trained with complementary objectives: inferring the action from an observed pre/post transition, predicting the post-condition from the pre-condition and action, and reconstructing the pre-condition from the post-condition and action. Across controlled ablations, visual transition simulation improves grounded word acquisition when it is paired with sufficient language supervision. The gains are concentrated on verbs rather than nouns, and within verbs, on result verbs rather than manner verbs, consistent with the fact that result verbs are defined by observable state changes while manner verbs depend more on motion trajectories. These findings suggest that grounding verb semantics requires more than aligning words with percepts: it requires predictive models of how the world changes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.