acceptodds
Under review as a conference paper at ICLR 2027

Learning Temporally Contextualized Latent Action Space for Robot Manipulation

Abstract

Robot policies are commonly trained to predict low-level action chunks from observations. However, raw action spaces are weakly aligned with task progress and lack inherent temporal semantics. In this work, we propose an action representation learning framework that learns a Temporally Contextualized Latent Action (TCLA) space. Beyond action reconstruction, we introduce two pretraining objectives: a temporal forward prediction objective in latent space and a task-progress estimation objective using episode-level spectral action context. Through experiments, we show that our method consistently improves training efficiency, success rates, and motion smoothness across pretrained Vision-Language-Action (VLA) models and World Action Models (WAMs), as well as policies trained from scratch. For example, on a RoboTwin task set, TCLA improves X-VLA from 55.6% to 89.0% in just 10k steps, and improves FastWAM from 20.6% to 40.8% at epoch 5. It also provides a new perspective on extending single-task policies to multi-task settings, highlighting the broader potential of temporal action pretraining.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.