CPR: Reviving Hard-Task Learning in Multi-Task Reinforcement Learning
Abstract
Multi-task reinforcement learning (MTRL) aims to learn multiple tasks with a shared model, yet hard tasks often make limited progress when trained alongside easier ones. We identify weak temporal-difference (TD) learning signals as a key characteristic of this failure and introduce expectile temporal-difference error (XTD), an upper expectile of signed TD errors that reveals informative positive value corrections obscured by averaging. Building on XTD, we propose Cross-level Progress Recovery (CPR), which revives hard-task learning through complementary inter-task and intra-task mechanisms. Inter-task dropout uses relative return and declining task-level XTD to identify sufficiently learned tasks and mitigate shared-critic over-specialization, while intra-task prioritization emphasizes high-XTD states to strengthen informative learning signals within hard tasks. Experiments on Meta-World show that CPR substantially improves overall performance, particularly on hard tasks, without relying on parameter resets. Trajectory and component analyses show that XTD highlights critical states and support the complementary roles of both mechanisms in reducing critic interference and improving hard-task learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.