Freeze the Dynamics, Learn the Interface: Modular Recurrent Scaffolds for Compositional World Models
Abstract
World models often learn representations and transition dynamics jointly, entangling perception with task structure and limiting compositional generalization. We introduce a hippocampal-inspired world model built on pretrained recurrent modules whose dynamics are frozen during downstream learning. Their multiscale periodic codes jointly disambiguate states that would be aliased within any individual module, providing a stable and structured latent coordinate system. A lightweight learned interface binds observations and actions to this scaffold, allowing the model to recover the topology of new task spaces with substantially fewer trainable parameters than end-to-end alternatives. The resulting representation decomposes task-relevant factors across modular dynamical subspaces and supports their recombination in previously unseen configurations. By exploiting this learned task-space structure, the model improves out-of-distribution compositional generalization and maintains more accurate long-horizon rollouts. The same structured latent geometry also enables goal-conditioned navigation and planning from state differences. These results suggest that pretrained multiscale grid-like dynamics can serve as a reusable scaffold for learning structured, compositional world models efficiently.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.