World Models as Version Spaces: Safe Tool Use under Hidden Rules
Abstract
World models enable tool-using agents to anticipate action outcomes and plan before execution. However, identical interfaces and similar observations can conceal different execution rules, making a plan safe in one deployment and unsafe in another. Overlooking this ambiguity can lead to irreversible errors that later replanning cannot undo. To address this challenge, we introduce VERSA, a world model that represents a version space of execution rules that remain plausible given the available evidence. VERSA learns a joint rule posterior through conditional flow matching and predicts plan outcomes with a learned rule-conditioned transition model. VERSA uses these predictions to select a plan under a risk constraint and decide whether further information is worth acquiring before execution. Theoretically, we derive upper bounds on execution risk and the value of further information, characterizing conditions under which near-optimal decisions can be made without fully identifying the deployment rules. Empirical evaluations show that VERSA achieves superior task performance and execution safety in hidden-rule environments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.