AWM: Accumulative World Modeling for Continual Reinforcement Learning
Abstract
Continual model-based reinforcement learning requires agents to adapt to non-stationary environments while preserving predictive dynamics from earlier tasks. A fundamental challenge lies in balancing the plasticity required to capture diverse, non-stationary dynamics against the stability needed to shield historical representations from backward interference. To resolve this fundamental problem, we introduce Accumulative World Modeling (AWM), a continual world-modeling framework that separates the capacity required to acquire new tasks from that needed to preserve past dynamics, while maintaining an evolving, trainable shared world model. Specifically, Composable Refinement Mechanisms (CRM) seamlessly adapt the shared world model by composing newly acquired adapters with those from prior tasks. Return-Constrained Compression (RCC) then compresses the adapters into a compact set of atoms, preventing parameter explosion. For task-agnostic inference, Confidence-Guided Expert Routing (CGER) uses cumulative observation-reconstruction error to identify the appropriate expert without task identities. Across two public benchmarks, AWM outperforms the strongest baselines, improving normalized final average performance (ACC) by 212.3% on Atari and 15.5% on CoinRun. Code will be publicly available upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.