acceptodds
Under review as a conference paper at ICLR 2027

Learning Skills from Historical Action Trajectories: Action Experience Dictionary for World Action Models

Abstract

World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experience across manipulation tasks. Furthermore, existing WAMs struggle to capture underlying cross-task semantic relationships that could guide target action prediction, as redundant background elements interfere with the extraction of key visual information. To address these challenges, we develop a novel Action Experience Dictionary (AED) that encodes historical physical action trajectories into shared action embeddings to support skill reuse and model cross-task relationships. Specifically, we first aggregate historical actions to align with visual observations and retrieve action embeddings from the AED using a pretrained action tokenizer. Subsequently, we visually condition the pooled embeddings through cross-attention and prepend them to noisy action tokens, providing interaction context and action intent for prediction. To encourage action-related motion modeling and reduce reliance on irrelevant background cues, we introduce a motion-aware transition loss that supervises visual feature change prediction over random temporal intervals. Experiments on simulation benchmarks and in real-world cross-embodiment settings demonstrate that our model outperforms the baselines. The anonymous project website is available at [https://anonymous.4open.science/w/supplementary_materials-4586/](https://anonymous.4open.science/w/supplementary_materials-4586/).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.