acceptodds
Under review as a conference paper at ICLR 2027

Identifying What Actions Read and Change in World Models

Abstract

World models support planning by predicting action consequences. Yet accurate prediction can leave latent state variables mixed across learned variables, obscuring which state variables the behavior policy reads (the read set) and which its actions change (the effect set). We jointly model how the behavior policy generates actions and how actions enter the transition, allowing next-step variables to depend on one another. We show that, under explicit conditions, video and logged actions reveal both roles of each variable that within-step dependence sets apart: its learned counterpart carries the same two labels, even while others stay mixed. The guarantee holds for models that reproduce the logged frame–action process with the fewest conditional dependencies among next-step variables. Under the same conditions, when next-step variables are conditionally independent given the current state and action, the guarantee covers every dynamic state variable. We learn one world model and estimate both sets on its frozen representation. Controlled latent experiments show closer variable matching and better read-set recovery than the same model trained for prediction alone, at comparable predic- tive accuracy. In a controlled RGB study on a fixed representation, read and effect estimates track policy and actuator switches, respectively. The selected variables also predict action responses more accurately than the full representation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.