acceptodds
Under review as a conference paper at ICLR 2027

TatRL: Closed-Loop World Model Evolution for Long-Horizon LLM Planning

Abstract

Large language models learn implicit world models through next-token prediction, yet leveraging this capacity for long-horizon planning remains challenging. Single-step methods, including supervised fine-tuning and turn-level optimization, cannot capture this structure, as they implicitly assume that the quality of each action is independent of the trajectory that generated it. This assumption holds for single-step tasks such as mathematical reasoning, but breaks down in long-horizon sequential tasks where actions are optimal only as a consequence of preceding moves. To address this, we introduce closed-loop world model evolution, a framework that treats entire occupancy trajectories as the unit of both generation and evaluation. Building on the connection between trajectory-level evaluation and multi-scale world modeling, we develop a progressive temporal scaling algorithm that trains within a self-play framework at increasingly long horizons, analogous to varying the planning horizon in geometric horizon models. We evaluate our approach on eight complex long-horizon tasks from the DouDizhu, GuanDan, and poker domains and find that trajectory-level evaluation is both necessary and sufficient for learning compositionally robust policies, achieving significant improvements across all benchmarks. Moreover, we achieve 99.2% action-sequence consistency while using 20 less demonstration data than single-step baselines, highlighting the sample efficiency of closed-loop world model evolution.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.