Long-Horizon Textual World Modeling through Structured Reasoning
Abstract
World models must predict how an environment evolves under sequences of actions, enabling agents to compare possible futures and reason about counterfactual actions before acting. Long-horizon prediction is commonly obtained by recursively applying a one-step transition model, but intermediate errors can compound over time. Multi-step dynamics models instead condition on a sequence of future actions and predict their consequences directly, but become harder to learn as horizon grows: the model must track interacting state changes across the trajectory, endpoint supervision provides weak credit assignment, and intermediate predictions can remain plausible while losing information needed for later states. We show that these challenges can be addressed by casting the internal evolution of a multi-step transition as structured reasoning over textual world states: reasoning over sparse state changes reduces the burden of state tracking, a predictive-gain objective rewards the learned state for improving over a matched predictor that conditions on raw history instead, and intermediate predictive rewards supervise each state along the trajectory. Because these intermediate states are explicit textual representations of the world, they provide semantically meaningful targets that can be inspected, scored, and corrected during training. Across ScienceWorld, Jericho, and CEO-Bench, our approach achieves the strongest average long-horizon performance against recursive and non-recursive baselines that condition directly on raw history, with gains increasing at longer horizons: at horizon 10, it exceeds the strongest such baseline by 12.3% relative on ScienceWorld and 9.3% relative on CEO-Bench, and error on Jericho drops by nearly 30%. In a controlled counterfactual study, our model is also the only one with statistically significant sensitivity to future actions (p<0.001), while baselines show null or anti-correlated sensitivity once ground-truth geometry is accounted for. These results suggest that making a multi-step transition's internal trajectory explicit and supervisable through structured reasoning is a better abstraction for long-horizon textual world modeling than recursively composing one-step predictions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.