Latent Action Policies from Suboptimal Data
Abstract
Latent Action Models (LAMs) learn control representations from action-free video, but existing pipelines typically train latent policies with behavioral cloning, without an explicit policy improvement objective. Although offline RL offers a natural approach to policy improvement, standard methods do not transfer directly to learned latent action spaces. We find that the structure of learned latent actions makes naive offline RL unreliable: directly maximizing value over latent actions or fitting a unimodal policy can produce latent actions that cannot be realized by physical execution, while the latent representation can merge physical controls that receive different rewards. To understand when latent-space optimization remains valid, we derive conditions under which policy improvement transfers to decoded execution and under which latent and physical optima align. Building on these insights, we propose LAPIS for goal-conditioned latent policy improvement, grounding value-based policy improvement in observed state–latent transitions and using a conditional flow policy to capture their multimodal structure. Across visual manipulation and navigation tasks, LAPIS consistently improves over latent-space baselines, while retaining strong performance on cube tasks even with only 1% action labels.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.