acceptodds
Under review as a conference paper at ICLR 2027

Evo-DexBench: Benchmarking Generalization across Related Tasks for Dexterous Manipulation

Abstract

General-purpose dexterous manipulation requires policies to transfer knowledge from related tasks to new tasks with only a few target demonstrations. Existing benchmarks often provide too many target demonstrations. They also leave the differences between source and target tasks uncontrolled. As a result, they cannot clearly distinguish knowledge transfer from task memorization. Our key insight is the Core-Gen principle. Core tasks are related tasks that share a functional interaction and provide source knowledge, while a Gen task is a novel target within the same functional family that has only a few demonstrations. This setup tests whether a policy can reuse functional interaction knowledge.Based on this principle, we propose Evo-DexBench, which controls both source-task relatedness and target-task exposure.It comprises 30 tasks in six functional families and a VR teleoperation pipeline. The resulting dataset covers three everyday scenes per task. It contains 1,248 multimodal demonstrations collected with simulated Shadow Hands equipped with 123 taxels. Evaluating representative policies, we find that the best achieves only 38.5% mean success, and that functionally matched Core data tends to yield larger gains than mismatched data. Results show that existing methods still struggle to extract and reuse functional interaction knowledge. Evo-DexBench provides an effective evaluation of the generalization ability of general-purpose dexterous manipulation policies across related tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.