World–Action Coupling and Collection-Matched Optimization for Online Multi-Task Reinforcement Learning
Abstract
Purely online multi-task reinforcement learning (MTRL) couples a global optimization clock to task-local data streams. With all-task updates at a fixed global frequency, task-wise update-to-data ratios can increase with the number of tasks. We introduce **WAM-MTRL-VE**, which couples latent future prediction to policy learning while matching task optimization to online data acquisition. The actor and an ensemble world model share a task-conditioned representation, and short imagined rollouts provide multi-step reward gradients to the actor through a frozen-model computation path. At each optimization tick, collection-matched stochastic optimization samples one replay-ready task according to the collection weights. Once all replay buffers are ready, matching update and collection probabilities recovers the nominal single-task update-to-data ratio in expectation while interleaving tasks. The entire agent is trained from scratch using online experience, without demonstrations or offline pretraining. On Meta-World, WAM-MTRL-VE reaches success on MT10 after 1M total interactions per run, averaged over five training seeds. On MT50, a checkpoint selected from one 12M-interaction training run achieves success across five independently seeded evaluation suites.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.