acceptodds
Under review as a conference paper at ICLR 2027

ReplayBind: Training Language World Models to Generate Action-Faithful Futures

Abstract

An agent deciding what to do next needs to distinguish the consequences of its alternatives. In language world models, however, the change that matters may occupy only a few words in an otherwise unchanged observation. Predicting that observation well can therefore leave the association between an action and its consequence weakly learned. We introduce ReplayBind, which learns this association from two observed branches sharing a visible state. Its crossed-likelihood objective favors the recorded action–outcome assignment over the assignment obtained by exchanging the outcomes, using forward prediction and inverse action reconstruction within one generator. Broad observation restoration and balanced branch replay carry this relational supervision into complete observation generation. Matched comparisons show improvements in action–outcome assignment, action reconstruction, and generated branch separation. These results show that ReplayBind better preserves action-induced changes across successive predictions, providing more reliable descriptions of future outcomes for distinguishing the consequences of different actions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.