Learning on-Policy Learning to Specialize Robot Policies
Abstract
We introduce Generative On-Policy Learners (GOLs), generative models that quickly adapt policy parameters to an unknown target domain. Conditioned on an initial (potentially random) policy and a few interaction episodes in the target domain, GOLs learn to generate improved policy parameters, by mimicking the policy evolution in traditional on-policy reinforcement learning. Unlike prior approaches, GOLs do not require (1) knowledge about domain parameters and rewards, (2) a large number of target domain interactions, or (3) expensive online domain identification. Generation is performed in a behavior-aligned latent space. We evaluate GOLs on continuous control problems, where they outperform state-of-the-art baselines in few shot domain adaptation on the most dynamic tasks. In addition, we demonstrate that our weight generation approach can iteratively improve policies in a sim-to-real transfer task. Finally, we characterize the key properties of the training dataset required to effectively train the weight generator, offering insights for scaling generative policy synthesis.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.