Value-Guided Variational Latent Dynamics for Offline Skill Stitching
Abstract
We study building agents capable of solving long-horizon tasks in the context of offline reinforcement learning (RL). Existing RL methods effectively learn individual skills. However, seamlessly combining these skills to tackle long-horizon tasks presents a significant challenge, as the termination state of one skill may be unsuitable for initiating the next skill, leading to cumulative distribution shifts. Previous works have studied skill stitching through online RL, while such methods can be time-consuming and may raise safety concerns when learning in the real world. Some other trajectory chaining approaches in previous literature can be utilized as offline skill stitching, but they either rely on time-intensive model-based planning, or neglect the underlying dynamics system implicit in the environment. In this paper, we propose to encode the states separated in time by several steps in a self-supervised manner via a variational auto-encoder. We then align the latent representations with the real actions, modeling the latent dynamics. Given that the aggregated datasets from all skills provide diverse and exploratory data, such latent dynamics can likely be leveraged to generate necessary transitions for stitching skills. During inference, we sample a population of latent actions and make selections according to a score-based state assessment. Extensive experimental results across benchmarks of both navigation and manipulation validate the effectiveness of our approach in comparison to baseline methods under offline settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.