Benchmarking and Enhancing Skill-Level Memory for Partially Observable Robotic Manipulation
Abstract
Visual robotic manipulation methods, including conventional VLA, WAM, etc., typically infer actions from the current observation or a limited temporal context. However, skill-level manipulation is often visually partially observable: task-relevant objects may become occluded, previously observed references may disappear from the current scene, and visually similar observations may correspond to different stages of skill execution. We introduce HIDE, an RLBench-based benchmark for evaluating skill-level memory under geometric, referential, and procedural partial observability. It contains 15 tasks with procedurally generated initial configurations, in which the correct action cannot always be determined from the current observation alone. We further investigate three memory mechanisms that respectively preserve task-relevant spatial information, historical references, and skill execution progress. Experiments reveal that these mechanisms exhibit distinct capability profiles: each provides clear benefits along its targeted dimension, while potentially causing limited degradation or interference on other dimensions. Their combination achieves the strongest overall performance by exploiting their complementary strengths. These findings demonstrate that different forms of visual partial observability require distinct memory mechanisms, and that maintaining internal representations of latent task states is essential for reliable skill-level robotic manipulation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.