Associative Dreamer: Predicting Action Consequences with Associative Memory
Abstract
In partially observed environments, identical current images and actions can have different consequences because of earlier events. We study how a world model can recover and use the relevant historical evidence within a bounded online state. We introduce Associative Dreamer, which learns to write visual associations into a fixed-size fast-weight memory, retrieve evidence through autoregressive queries, and condition its action-dependent latent prior on that evidence. At known retrieval times, an auxiliary objective aligns retrieved representations with exponential moving average (EMA) targets of annotated, query-relevant past observations. Targets come from the same training run and serve only as loss supervision. We construct a benchmark of three controlled pixel tasks whose future action consequences depend on earlier evidence, and evaluate short open-loop prediction on fixed training trajectories. Within a 20,000-update budget, the aligned model achieves 99.58% and 100% mean final-window consequence accuracy in Transport and Power, compared with approximately 50% for both an ordinary recurrent model and memory trained without alignment. At 40,000 updates in Assembly, aligned and prediction-only memory average 77.01% and 75.97% balanced accuracy, respectively. Disabling future reads removes the successful models' predictive advantage. These results show that Associative Dreamer combines retrieved historical evidence with the current state and actions to predict future consequences that depend on earlier experience.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.