FlexMemWAM: Flexible Temporal Compression for World Action Models
Abstract
Long-horizon robotic manipulation requires world action models (WAM) to retain historical evidence while controlling memory costs. Task-relevant changes are unevenly distributed over time: many consecutive observations are redundant, whereas brief interactions can introduce important new evidence. Keyframe selection may discard such evidence, while fixed per-frame compression assigns the same memory capacity to every observation. We propose FlexMemWAM, which uses online robot-state cues to organize history into variable-length segments, each sharing a fixed set of learnable slots. Once a boundary is causally confirmed, a shared set of learnable slots jointly reads the segment and prior context within the WAM’s video Transformer layers, producing layer-wise key–value memories for subsequent action prediction. The memory slots are optimized through joint video and action prediction. FlexMemWAM achieves average success rates of 84.1% on nine history-dependent RMBench tasks and 71.7% on three real-world memory-required tasks, outperforming other strong competitors. We also show that FlexMemWAM effectively reduces the historical KV story compared with per-frame allocation. The code will be made publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.