acceptodds
Under review as a conference paper at ICLR 2027

Can We Check Whether a World Model Follows Actions Without a Simulator?

Abstract

World models predict how an environment changes when an agent acts, and agents plan by comparing these predictions across actions. This only works if changing the action moves the prediction the way the real environment would move. Many existing evaluations check this with cheap probes that need no simulator. We test whether these probes notice when a model’s prediction moves the wrong way. We use simulators that can return to a saved state and run again under a different action. This gives the true effect of each action change. We study 92 robot world models that we trained and 143 released models, including models that predict Atari frames. We also reflect each model’s response to actions, which reverses its direction and keeps its size. We prove that probes that look only at the size of the response cannot detect this, and in our experiments they do not. A popular probe, logged action retrieval, even scores the reflected models higher. Training models on its criterion improves their retrieval score while their true direction gets worse. We also propose the matched probe, which compares the model’s response with pairs of similar situations in the recorded data. It tracks the true direction in all three groups of models. We recommend testing any action probe on reflected models before trusting it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.