Capturing Inter-Task Reinforcement Learning Dynamics for Adaptive Task-Sampling
Abstract
Multi-Task Reinforcement Learning (MTRL) trains a single agent on multiple tasks simultaneously, potentially improving performance and sample efficiency through positive transfer. However, MTRL suffers from *imbalanced learning*: tasks progress at different rates and may interfere with one another. Prior work addresses this problem using specialized task-sampling strategies that prescribe, *a priori*, how experience should be allocated across tasks. We instead argue that task-sampling strategies should be learned directly from data. For that end, we introduce DynATS, a learned sampler that captures inter-task learning **D**ynamics for **A**daptive **T**ask **S**ampling. DynATS is trained across sets of tasks using Evolution Strategies to maximize the performance of the multi-task agent as a whole. On MTRL problems, DynATS outperforms existing sampling strategies on average. By construction, DynATS transfers seamlessly across task distributions and domains. We transfer a learned sampler trained only on discrete, grid-based task sets to unseen task sets from the same distribution and to unseen continuous-control tasks in MuJoCo and Meta-World, demonstrating that the learned strategies can generalize beyond the environments on which they were trained.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.