acceptodds
Under review as a conference paper at ICLR 2027

DOLA: Latent Action Pretraining from Delta Observations

Abstract

Latent action models learn action representations from action-free video, scaling visuomotor learning beyond the limited labeled demonstrations available in robotics. These models typically train a latent inverse dynamics model jointly with a forward dynamics model on paired observations, hoping the latent captures the underlying action through bottlenecked reconstruction. However, since actions cause state changes, an action representation should reflect the change in state it produces. Building on this inductive prior, we present DOLA (Delta-Obs Latent Actions), a new latent action model that explicitly conditions the inverse dynamics on the difference between consecutive observation embeddings, jointly trained with the encoder and an autoregressive forward dynamics model that reconstructs multiple future embeddings. On a suite of four simulated environments, DOLA learns latent actions that linearly recover ground-truth actions across both manipulation and navigation tasks. A single DOLA model trained on unpaired data from two robots, grounded by heterogenous joint velocity actions, learns a shared latent space that linearly recovers end-effector Cartesian deltas and supports open-loop trajectory transfer across embodiments. On a real Franka arm, DOLA executes zero-shot video plans from an off-the-shelf video generator with 75 minutes of task-agnostic random calibration data. Finally, we ablate key components of DOLA and measure its impact on representation quality and downstream task performance. Videos are best viewed at https://dola-anon.github.io

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.