Keep Everything, Read One Latent: Verbatim Episode Memory for World-Action Models
Abstract
World-action models (WAMs) typically act from a bounded window of recent ob- servations, which makes their decisions effectively Markovian. We show that a WAM becomes history-dependent with verbatim episode memory (VEM). Every observation, from the demonstration and from execution alike, stays in context as its encoded video latent. After temporal sampling and native VAE encoding, no additional memory compression, selection, or eviction is applied. Training uses expert actions and future video targets, without memory annotations or subgoal supervision. The video expert integrates the causal history into the current latent, while the action expert reads only this fixed-size, history-conditioned state. VEM achieves state-of-the-art overall performance on RoboMME, with a 62.5% aver- age success rate across 16 tasks. Placing the demonstration and execution history on one stream elicits strong imitation capabilities, yielding a 78.2% success rate on the Imitation Suite. This outperforms policies conditioned on ground-truth sub- goals on both our backbone (49.5%) and a language-conditioned VLA (64.0%). In a retrieval test, a single initial observation still steers behavior after more than twice the length of the longest training episode of irrelevant history.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.