Compositional Context Encoding for Out-of-Distribution Generalization in Modular World Models
Abstract
World models used in model-based reinforcement learning must generalize to physical configurations unseen during training. A key challenge arises in compositional out-of-distribution (OOD) settings: the parameters of each mechanism span a wide range, but training varies each mechanism in isolation, whereas test time requires simultaneous extrapolation across all mechanisms. We show that a naive global context encoder fails in this setting because the joint OOD input it encounters at test time was never seen in any form during training, even though the parameters of each individual mechanism remain within the training distribution. We propose Compositional Context Encoding (CCE), which assigns one independent GRU encoder to each causal mechanism. Grounded in the Independence of Causal Mechanisms (ICM) principle, each encoder reads only the transition history of its own state subspace, so the joint OOD test input is automatically decomposed into per-encoder single-sided OOD inputs that the training distribution does cover. A lightweight auxiliary loss aligns inferred mechanism parameters with oracle values during training, enabling stable long-horizon rollouts. On a decoupled double-pendulum benchmark with 64 compositional OOD test configurations, CCE achieves – lower 20-step rollout MSE than a parameter-free monolithic baseline across three random seeds. By contrast, a global-encoder variant performs worse than the parameter-free baseline, confirming that the compositional structure is essential.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.