MOTO-WM: Learning Multiple Tasks Online with One World Model
Abstract
World models support model-based reinforcement learning through planning and imagined policy optimization. Extending them to multiple tasks requires sharing predictive structure while preserving differences between task objectives. Through empirical and representation analyses of independently trained DreamerV3 agents, we find shared physical structure in recurrent states but objective-dependent reward and policy readouts. Motivated by this observation, we propose Multi-task Online Training with One World Model (MOTO-WM), which shares a world model and actor–critic across tasks while conditioning feedback and control on task identity. A unified action interface supports different action dimensions. MOTO-WM learns online from scratch without demonstrations, pre-collected data, or pretrained models. On 20 visual DeepMind Control tasks, a single 10.55M-parameter model outperforms the evaluated model-based baselines at 500K steps per task with only 1.48 GPU-days of training. Meta-World results, ablations, and controlled task interventions further validate the method and its task-conditioned behavior. Essential scripts and checkpoint are provided in the supplementary material.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.