Evo-DexBench: Benchmarking Generalization across Related Tasks for Dexterous Manipulation
Abstract
General-purpose dexterous manipulation requires policies to transfer knowledge from related tasks to new tasks with only a few target demonstrations. Existing benchmarks often provide too many target demonstrations. They also leave the differences between source and target tasks uncontrolled. As a result, they cannot clearly distinguish knowledge transfer from task memorization. Our key insight is the Core-Gen principle. Core tasks are related tasks that share a functional interaction and provide source knowledge, while a Gen task is a novel target within the same functional family that has only a few demonstrations. This setup tests whether a policy can reuse functional interaction knowledge.Based on this principle, we propose Evo-DexBench, which controls both source-task relatedness and target-task exposure.It comprises 30 tasks in six functional families and a VR teleoperation pipeline. The resulting dataset covers three everyday scenes per task. It contains 1,248 multimodal demonstrations collected with simulated Shadow Hands equipped with 123 taxels. Evaluating representative policies, we find that the best achieves only 38.5% mean success, and that functionally matched Core data tends to yield larger gains than mismatched data. Results show that existing methods still struggle to extract and reuse functional interaction knowledge. Evo-DexBench provides an effective evaluation of the generalization ability of general-purpose dexterous manipulation policies across related tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.