Selective Recall for Late-Bound Robotic Manipulation
Abstract
People encounter a stream of events throughout the day, yet a later request—such as fetching an object put away hours earlier—may depend on only a few relevant memories. Inspired by this everyday need for selective recall, we study how robots can use accumulated experience when the task is not known in advance. Existing manipulation-memory evaluations often reveal goals during history encoding, place key evidence at fixed temporal positions, or associate each history with a single task, leaving task-conditioned access to rich, task-agnostic experience insufficiently evaluated. We introduce late-bound manipulation: a robot first experiences a long, continuous history containing numerous candidate events interleaved with background interactions, and only a subsequent instruction determines which evidence is needed for action. We develop MOSAIC, which pairs identical histories and execution-start states with different instructions requiring evidence from different temporal intervals. Its shared-history construction enables controlled comparisons of behavior under different subsequent instructions. We further propose ReWAM, a selective-memory world-action model that combines compact global memory, instruction-informed retrieval of detailed historical chunks, and recent local context for video and action prediction. On seven RMBench tasks, ReWAM achieves 50.57% mean success, compared with 55.00% for dense-history conditioning and 44.71% for dense visual history with local action context. On the evaluated three- and six-target MOSAIC configurations, it achieves 93% and 45% success, compared with 36% and 26% for dense-history conditioning. In separate timing experiments at 256 historical video chunks, ReWAM achieves a 3.06× full-cycle inference speedup over dense-history conditioning, with latency increasing by only 4.27% from 1 to 256 chunks. Together, MOSAIC and ReWAM provide a framework for studying task-conditioned recall and its performance–cost trade-offs in long-history manipulation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.