acceptodds
Under review as a conference paper at ICLR 2027

Thinking by Doing: Building Efficient World Model Reasoning in LLMs via Multi-turn Interaction

Abstract

A central challenge for LLM agents is modeling how environment states evolve under actions, which calls for robust world modeling to improve foresight and enhance decision-making under uncertainty. Multi-turn RL provides direct feedback for learning these transitional dynamics, yet current approaches impose fixed reasoning structures that limit models to surface-level heuristics, often at the cost of inefficient interaction or over-dependence on external signals. Whether these interactive experiences truly lead to internalized knowledge of state transitions, however, remains unclear. To this end, we explore world-model internalization through efficient interaction and active reasoning (WMAct), which uses two mechanisms to convert interaction into internalized transition knowledge without prescribing a reasoning template: (1) reward rescaling that weights the outcome reward by the proportion of state-changing actions, discouraging ineffective interaction; (2) interaction-frequency annealing that gradually tightens the turn budget, encouraging less reliance on online feedback and more on internal transition prediction. Across Sokoban, Maze, and Taxi, our results reveal that \methodabb enables effective single-turn solving of tasks that previously demanded multi-turn interaction. This performance gain, achieved without recourse to multi-turn feedback, is indicative of genuine planning from internalized state transitions rather than reactive pattern matching. Moreover, the learned capability transfers to general benchmarks (e.g., +5.05 on HMMT25, +2.24 on GPQA-Diamond).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.