acceptodds
Under review as a conference paper at ICLR 2027

Learning on-Policy Learning to Specialize Robot Policies

Abstract

We introduce Generative On-Policy Learners (GOLs), generative models that quickly adapt policy parameters to an unknown target domain. Conditioned on an initial (potentially random) policy and a few interaction episodes in the target domain, GOLs learn to generate improved policy parameters, by mimicking the policy evolution in traditional on-policy reinforcement learning. Unlike prior approaches, GOLs do not require (1) knowledge about domain parameters and rewards, (2) a large number of target domain interactions, or (3) expensive online domain identification. Generation is performed in a behavior-aligned latent space. We evaluate GOLs on continuous control problems, where they outperform state-of-the-art baselines in few shot domain adaptation on the most dynamic tasks. In addition, we demonstrate that our weight generation approach can iteratively improve policies in a sim-to-real transfer task. Finally, we characterize the key properties of the training dataset required to effectively train the weight generator, offering insights for scaling generative policy synthesis.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.