Hierarchical Reinforcement Learning with Temporally-extended Latent Action World Models
Abstract
World Models enable counterfactual reasoning and planning but struggle in long horizon tasks because prediction errors compound over the horizon. Temporal abstraction can mitigate these errors by simply requiring fewer consecutive predictions for a given horizon. We propose a Temporally-extended Latent Action World Model (Te-LAWM) and present a hierarchical model-based framework for goal-conditioned reinforcement learning that leverages it to generate subgoals for a value-based model-free policy. Our temporally-extended latent action model learns a multi-step action abstraction, a multi-step dynamics model conditioned on these abstract actions (the Te-LAWM), and a prior policy over the actions within a single objective. Subgoal planning is done by rolling out the prior policy within the Te-LAWM. We demonstrate superior performance of our framework compared to flat model-based and hierarchical model-free approaches on benchmark navigation and manipulation tasks. Ablations show that autoregressive subgoal generation allows the agent to stitch together experience from different trajectories, and that temporal action abstraction is essential for scaling world model predictions to long horizons and large action spaces.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.