acceptodds
Under review as a conference paper at ICLR 2027

Rules Are Rare: When World Models Memorise Rather Than Acquire Rules

Abstract

World models are trained to predict the next state, and by that objective they also learn the rules of their environment. We show that the world models we train learn them by memorising. On Craftax, a transformer world model (TWM) is as accurate on a rule before the episode has shown it as after (88% against 89%), and when the rules are resampled each episode it falls to the accuracy of predicting that nothing changes. Training on changing rules does not repair this: on XLand-MiniGrid, a Long-Context Transformer model that sees the whole episode memorises one rule table per training episode instead; updating the weights does, by fine-tuning, replay, hidden-parameter inference or online test-time adaptation (AdaJEPA), at three separable costs: interference, forgetting and erosion. The hidden parameter in these environments is not one vector: it splits by interaction type, and the interaction is observable while its effect is not. So we introduce PRISMO. It keeps one entry per interaction type in a store outside the frozen host, writes it from a single observation with no gradient step, and clears it when the rules may change. On two frozen Craftax hosts it raises rule-state accuracy from 29% to 87% (TWM) and from 24% to 79% (a long-context Transformer) while every other predicted token is unchanged, and on XLand two rules observed once each compose into a chain never seen executed.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.