How a World Model Forms: Separating World Knowledge from Behavioral Competence in Sokoban Transformers
Abstract
Does teaching a policy to represent its environment more accurately make it better at acting? We study this question in Sokoban Transformers that receive an initial board and move history, without updated observations. Eight training conditions vary supervision of state, action effects, and goal occupancy while sharing the same action targets and trajectory corpus. We measure what probes recover, which state distinctions guide decisions, and whether the policy solves. Action prediction alone yields accurate action-effect readouts and decisions that distinguish remote states, despite incomplete state reconstruction. Across three seeds, supervising all three components raises complete-state recovery from 0.208 to 0.955, while solving on held-out puzzles with optimal distances of 32–60 changes from 0.745 to 0.735. Supervision does help when policies must continue from supplied detour histories, but doubling training in both conditions reduces this advantage, leaving training efficiency as a plausible contributor. Together, these results show why complete-state recovery cannot stand in for behaviorally useful world knowledge. Its value must be tested through the decisions and continuations it is intended to support.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.