acceptodds
Under review as a conference paper at ICLR 2027

Double Take: Intrinsically Motivating Agents to Model Many Worlds

Abstract

We consider the problem of an agent with prior beliefs over a distribution of possible worlds, who wants to infer the identity of the new world in which it finds itself at runtime. The key is to take the right kind of exploratory actions. Supposing the agent has some implicit belief about the current world, we show that a surrogate for information gain can be computed by training a model to predict multiple possible outcomes of an action given some observed history in the world, then asking the model how much more likely an outcome becomes after observing it once. The benefit of such a "double take" is that first looks are informative only insofar as they reveal the hidden structure of the environment; by construction, they cannot help predict noisy TVs. However, assuming a black-box environment, we cannot sample alternate outcomes conditioned only on an agent's observations if the environment includes noisy hidden state. To get around this, we instead perform an entire "second take" of agent interaction in the same world, and train a model to better predict the second take using information from the first. This teaches the model to transmit shared-world structure and discard rollout-specific transients. We can then simulate the effect of an instantaneous double take by replaying the first look to the model as a transition within a supposed first take, and having it make new predictions using the shared-world structure this disguised observation confers. We use procedurally-generated reinforcement learning environments to model the problem of information gathering over many worlds, measuring the amount of information gathered about the world by the resultant agents using an ideal Bayesian observer. Our methods outperform intrinsic-RL baselines in terms of world-information-gathering, and often discover behavior correlating with high environment reward. They may also provide a general recipe for collecting relevant data for training interactive next-frame-prediction world models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.