acceptodds
Under review as a conference paper at ICLR 2027

Off the Record: What a World Model Guarantees at Unrecorded Actions

Abstract

World models let an agent plan by asking what an action would do instead of taking it. They are trained and validated on transitions recorded under a behavior policy, yet planning asks about actions that policy never took. We ask what a world model can guarantee at such actions. It is elementary that the data alone guarantees nothing there, since a model can fit it perfectly and still be arbitrarily wrong. Our two results say what a guarantee rests on. First, a Lipschitz constant declared for the model's predictor turns distance from the data into a bound. The bound takes two assumptions: the true dynamics respect that constant, and the model's encoder retains the relevant state. Under them, the next-state error of every model that fits the data and respects the constant is bounded from above by a band that widens with that distance. The band is attained in the worst case, and a constant declared too low can make it silently false. Half the disagreement between two models bounds the larger of their errors from below, with no constant. Second, the encoder sets that distance, and discarding nuisance shrinks it. On MuJoCo datasets, models that score alike on held-out data diverge at unrecorded actions, and their error and disagreement grow with that distance. An adversarial search finds an error within a small factor of the band. Replacing part of a dataset with widely varied actions lowers error far from the data. We turn these results into guiding principles. Distance from the data, not a held-out score, should govern how far a world model's prediction is trusted.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.